> ## Documentation Index
> Fetch the complete documentation index at: https://docs.abliteration.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate a dataset

Use these complete Python examples to generate records and save a JSONL dataset.

## Set up

You need Python, a [Project API key](https://abliteration.ai/console/api-keys), and API credits.

```sh Install SDK wrap theme={"system"}
python -m pip install openai
```

Set your API key in the terminal:

<CodeGroup>
  ```sh macOS / Linux wrap theme={"system"}
  export ABLIT_KEY="YOUR_API_KEY"
  ```

  ```powershell Windows PowerShell wrap theme={"system"}
  $env:ABLIT_KEY = "YOUR_API_KEY"
  ```
</CodeGroup>

## Choose settings

Replace the values in your chosen Python example:

| Python argument | Typical effect |
| - | - |
| `temperature=0.2` | More consistent wording; useful for focused question-and-answer examples. |
| `temperature=0.7` | More variation in wording; a starting point for diverse support messages. |
| `temperature=1.0` | More randomness; try it for varied scenarios and review the output carefully. |

Temperature changes sampling, not factual accuracy. Even low-temperature runs can produce different results.

Choose `model="abliterated-model"` (Base) or `model="abliterated-model-large-v2"` (Large V2) in the examples below.

Use `reasoning_effort="low"` to start. Try `"high"` for more reasoning; it can take longer and use more tokens.

## Choose an example

Expand one example and copy the complete script into `generate.py`. Edit `PROMPT` to describe your dataset.

<Tabs>
  <Tab title="Support">
    ```python generate.py expandable wrap theme={"system"}
    import json
    import os
    from openai import OpenAI

    PROMPT = (
        'Generate 10 fictional customer support messages. '
        'Return JSON with a records array. '
        'Each record needs text and label. '
        'Labels: billing, account_access, technical_support. '
    )

    client = OpenAI(
        base_url="https://api.abliteration.ai/v1",
        api_key=os.environ["ABLIT_KEY"],
    )

    response = client.chat.completions.create(
        model="abliterated-model",
        temperature=0.7,
        reasoning_effort="low",
        max_tokens=4096,
        response_format={"type": "json_object"},
        messages=[{
            "role": "user",
            "content": PROMPT,
        }],
        extra_body={"include_reasoning": False},
    )

    choice = response.choices[0]
    if choice.finish_reason != "stop":
        raise RuntimeError(
            "Incomplete output. Increase max_tokens or request fewer records."
        )

    records = json.loads(choice.message.content)["records"]
    if not isinstance(records, list) or not all(
        isinstance(row, dict) for row in records
    ):
        raise ValueError("Expected a records array of JSON objects.")
    with open("dataset.jsonl", "x", encoding="utf-8") as output:
        for record in records:
            output.write(json.dumps(record, ensure_ascii=False) + "\n")

    print(f"Saved {len(records)} records to dataset.jsonl")
    ```
  </Tab>

  <Tab title="Q&A pairs">
    ```python generate.py expandable wrap theme={"system"}
    import json
    import os
    from openai import OpenAI

    PROMPT = (
        'Create 10 question and answer pairs '
        'about Python virtual environments. '
        'Return JSON with a records array. '
        'Each record needs question and answer. '
    )

    client = OpenAI(
        base_url="https://api.abliteration.ai/v1",
        api_key=os.environ["ABLIT_KEY"],
    )

    response = client.chat.completions.create(
        model="abliterated-model-large-v2",
        temperature=0.5,
        reasoning_effort="low",
        max_tokens=8192,
        response_format={"type": "json_object"},
        messages=[{
            "role": "user",
            "content": PROMPT,
        }],
        extra_body={"include_reasoning": False},
    )

    choice = response.choices[0]
    if choice.finish_reason != "stop":
        raise RuntimeError(
            "Incomplete output. Increase max_tokens or request fewer records."
        )

    records = json.loads(choice.message.content)["records"]
    if not isinstance(records, list) or not all(
        isinstance(row, dict) for row in records
    ):
        raise ValueError("Expected a records array of JSON objects.")
    with open("dataset.jsonl", "x", encoding="utf-8") as output:
        for record in records:
            output.write(json.dumps(record, ensure_ascii=False) + "\n")

    print(f"Saved {len(records)} records to dataset.jsonl")
    ```
  </Tab>

  <Tab title="Web search">
    Enable [web search](/capabilities/web-search) in the Console first.

    ```python generate.py expandable wrap theme={"system"}
    import json
    import os
    from openai import OpenAI

    PROMPT = (
        'Use web search and official Python documentation '
        'to create 5 question and answer pairs '
        'about virtual environments. '
        'Return JSON with a records array. '
        'Each record needs question and answer. '
    )

    client = OpenAI(
        base_url="https://api.abliteration.ai/v1",
        api_key=os.environ["ABLIT_KEY"],
    )

    response = client.chat.completions.create(
        model="abliterated-model",
        temperature=0.4,
        reasoning_effort="low",
        max_tokens=8192,
        response_format={"type": "json_object"},
        messages=[{
            "role": "user",
            "content": PROMPT,
        }],
        extra_body={
            "include_reasoning": False,
            "web_search_options": {"search_context_size": "medium"},
        },
    )

    choice = response.choices[0]
    if choice.finish_reason != "stop":
        raise RuntimeError(
            "Incomplete output. Increase max_tokens or request fewer records."
        )

    records = json.loads(choice.message.content)["records"]
    if not isinstance(records, list) or not all(
        isinstance(row, dict) for row in records
    ):
        raise ValueError("Expected a records array of JSON objects.")
    sources = [a.url_citation.url for a in choice.message.annotations or []
               if a.type == "url_citation" and a.url_citation]
    with open("dataset.jsonl", "x", encoding="utf-8") as output:
        for record in records:
            record["response_sources"] = sources
            output.write(json.dumps(record, ensure_ascii=False) + "\n")

    print(f"Saved {len(records)} records to dataset.jsonl")
    ```
  </Tab>
</Tabs>

## Run it

```sh Run generator wrap theme={"system"}
python generate.py
```

Open `dataset.jsonl` in a text editor. Each line is one record. Choose a new filename for another run to keep the previous dataset. Use `python3` if that is your computer's Python command.

Review the records before using them. The web-search example saves `response_sources` for the whole response; check that the sources support each answer.

## Output

A support-message record looks like this. Each JSONL record occupies one line in the file:

```json dataset.jsonl wrap theme={"system"}
{"text":"How do I update my payment method?","label":"billing"}
```

For larger datasets, use [batch generation](/data-generation/labeled-datasets).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.