> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bagofwords.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create a demo dataset

> Ask the agent for realistic mock data for any domain, review the proposed schema and agents, and get a ready-to-use connection with agents on top.

In a Training session, you can ask the agent for a realistic, fictional database for any domain: finance, sales, HR, application logs, monitoring metrics, machine data and more. The agent designs the tables and suggests agents to build on them. You review the proposal, choose which agents to create, and the data is generated and connected for you.

Use it to demo Bag of Words to a team before connecting real data, to prototype agents and instructions, or to practice training and evals on data that looks like yours.

## Before you start

* **Demo data generation is on.** An admin turns it on in **Settings → AI Settings → Agent capabilities**, under **Training Mode**. It is off by default.
* **Training Mode is on.** It is on by default, in the same section.
* **You can manage connections.** The dataset is added as a new connection, so you need the permission to manage connections. Creating the suggested agents also needs the permission to create agents. Without it, only the connection is created.
* **A small default model is set.** The rows are generated by the organization's small default model. See [LLM settings](/using-bow/llm).

You don't need any agents or data connected first. A new workspace can start here.

## Step 1: Ask for the dataset

Start a Training session: open a new report and switch the prompt box from **Chat** to **Training**. Then describe the business and what you want to analyze:

> *I want a demo for finance in the e-commerce sector: orders, payments, refunds, chargebacks and payouts over the last 2 years. Build it and suggest agents.*

The agent picks a fictional company, designs the tables, writes how each column should behave (for example, "Q4 revenue is 40% higher" or "refunds only on delivered orders"), and proposes one to four agents.

## Step 2: Review the proposal

The proposal appears as a card titled **Review the demo dataset**:

* **Tables**: each table with what one row represents and how many rows it will have. Click a table to see its columns. A key icon marks the primary key and a link icon marks a column that references another table.
* **Suggested agents**: each agent with its icon, description and the tables it will use. Untick any agent you don't want.

When it looks right, click **Create dataset**. The button shows how many agents will be created.

To change something, click **Request changes**, describe what you want (for example, *add a budgets table so we can compare actuals vs plan*), and click **Send feedback**. Nothing is created. The agent revises the proposal and shows a new card.

<Note>
  The card waits about four minutes for your answer. If you don't respond in time, nothing is created and you can ask the agent to propose it again. If you reload the page while it waits, the card comes back.
</Note>

## Step 3: Watch it build

The card shows **Generating data…** and marks each table as it finishes. Tables that depend on others are generated after them, so every reference points at a real row, including tables that reference themselves, such as employees and their managers.

Each table is checked before it is saved: unique keys, references that resolve, the right column types, no empty values where they aren't allowed, and no dates in the future. If a table fails a check, it is regenerated automatically.

## Step 4: Use the result

When the card shows **Demo dataset created**, you have:

* **A new connection** marked **Demo · SQLite**, with a description for every table and column.
* **The agents you kept ticked**, each with its icon, its tables, conversation starters and instructions for its business definitions. Click **Open** to go to an agent.
* **The agents attached to this session**, so you can keep going right away.

Try a conversation starter, or keep training:

> *Inspect the data and check the instructions you wrote against it.*

> *Create eval cases for revenue and refund rate and run them.*

See [Train an agent](/agents/train-an-agent) and [Test an agent with evals](/guides/evals).

## Share the session

If you share the conversation, people with the link see the card read-only: **Waiting for review** while you review, then the dataset once it is created. The shared page updates on its own while the run is in progress.

## Prompts for other domains

| Domain | Prompt |
| - | - |
| HR | *Create a demo dataset for HR analytics at a 400-person software company: headcount, hiring, attrition and compensation. Suggest agents on top.* |
| B2B sales | *Build a demo dataset for a B2B SaaS sales team: accounts, contacts, opportunities through pipeline stages, activities and closed deals. Add suggested agents.* |
| Monitoring | *Create demo data for monitoring a microservices platform: services, hosts, request logs with severities, and per-minute latency and error metrics with a couple of incidents. Suggest agents.* |
| Manufacturing | *Create demo data for a factory: machines, sensor readings every 10 minutes, maintenance events and failures. Temperature and vibration should rise before a failure.* |

<Tip>
  The more you say about how the business behaves, the more realistic the data: seasonality, typical mixes ("70% card, 20% PayPal"), skew ("a few customers place most orders"), and events to include ("two outages in the last month").
</Tip>

## Limits

* Up to 15 tables and about one million rows per dataset.
* Up to 10 generated datasets per organization. Delete one to create another.
* Dates end today or earlier.
* All data is fictional: no real people, companies or valid identifiers such as card numbers.
* An agent name that is already taken gets a number, for example **Revenue & Refunds (2)**.

## Delete a demo dataset

Delete its connection. This deletes the generated database and any agent that uses only that connection.

<Note>
  Self-hosted: generated databases are stored in the backend's `uploads/demo_data/` folder. If you run more than one backend instance, or containers without persistent storage, put `uploads/` on a shared, persistent volume.
</Note>

## Related

* [Train an agent](/agents/train-an-agent)
* [Tips for training an agent](/guides/training-tips)
* [Test an agent with evals](/guides/evals)
* [AI Settings](/using-bow/settings)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.