> ## Documentation Index
> Fetch the complete documentation index at: https://docs.abliteration.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate a labeled dataset

Generate 500 fictional moderation records in batches and save them as JSONL.

## Set up

Complete the [Python setup](/data-generation/quickstart#set-up). You need a [Project API key](https://abliteration.ai/console/api-keys) and API credits.

## Choose settings

The complete script below includes these settings. Start with `COUNT = 10`, then increase it to `500`.

```python Settings wrap theme={"system"}
COUNT = 500
BATCH_SIZE = 10
LANGUAGE = "English"
REGION = "global"
MODEL = "abliterated-model-large-v2"
```

Use `abliterated-model` for general-purpose generation or Large V2 for more demanding text tasks. Both support `temperature` and `reasoning_effort` in the API request.

## Write the prompt

Edit `PROMPT` in the script to define your categories and policy criteria. This example requests:

| Field | Value |
| - | - |
| `category` | `privacy_exposure`, `threats`, `harassment`, or `benign` |
| `severity` | `low`, `medium`, or `high` |
| `language`, `region` | Your selected settings |
| `rationale` | A brief explanation of the label and severity |

## Generate

Expand and copy the complete script into `generate.py`. It includes batching, validation, duplicate checks, and JSONL saving.

<Accordion title="Complete generate.py script">
  ```python generate.py wrap theme={"system"}
  import json
  import os
  from openai import OpenAI

  COUNT = 500
  BATCH_SIZE = 10
  LANGUAGE = "English"
  REGION = "global"
  MODEL = "abliterated-model-large-v2"

  PROMPT = (
      "Create fictional content moderation records. "
      "Vary scenarios and wording. "
      "Copy these language and region values exactly. "
      "Return JSON with a records array. "
      "Fields: text, category, severity, language, region, rationale. "
      "Categories: privacy_exposure, threats, harassment, benign. "
      "Severity: low, medium, high. "
      "Low: benign or minor concern. Medium: targeted violation. "
      "High: credible risk of harm. "
      "Explain category and severity briefly in rationale. "
      "Use fictional names and placeholder contact details only. "
      "These are policy labels, not legal determinations."
  )

  client = OpenAI(
      base_url="https://api.abliteration.ai/v1",
      api_key=os.environ["ABLIT_KEY"],
  )

  saved = 0
  seen = set()
  with open("dataset.jsonl", "x", encoding="utf-8") as output:
      while saved < COUNT:
          size = min(BATCH_SIZE, COUNT - saved)
          response = client.chat.completions.create(
              model=MODEL,
              temperature=0.7,
              reasoning_effort="low",
              max_tokens=8192,
              response_format={"type": "json_object"},
              messages=[{
                  "role": "user",
                  "content": (
                      f"Generate {size} records. "
                      f"Batch {saved // BATCH_SIZE + 1}. "
                      f"Language: {LANGUAGE}. Region: {REGION}. "
                      + PROMPT
                  ),
              }],
              extra_body={"include_reasoning": False},
          )
          choice = response.choices[0]
          if choice.finish_reason != "stop":
              raise RuntimeError(
                  'Incomplete batch. Increase max_tokens or lower '
                  'BATCH_SIZE.'
              )
          records = json.loads(choice.message.content)["records"]
          if not isinstance(records, list) or len(records) != size:
              raise ValueError(
                  'Unexpected batch size. Revise the prompt and retry.'
              )

          batch_texts = set()
          for row in records:
              fields = (
                  "text", "category", "severity",
                  "language", "region", "rationale",
              )
              if not isinstance(row, dict) or any(
                  not isinstance(row.get(key), str) or not row[key].strip()
                  for key in fields
              ):
                  raise ValueError(
                      'A record is missing required text fields.'
                  )
              if row["category"] not in {
                  "privacy_exposure", "threats", "harassment", "benign",
              }:
                  raise ValueError(
                      'Unexpected category. Revise the prompt and retry.'
                  )
              if row["severity"] not in {"low", "medium", "high"}:
                  raise ValueError(
                      'Unexpected severity. Revise the prompt and retry.'
                  )
              if row["language"] != LANGUAGE or row["region"] != REGION:
                  raise ValueError(
                      'Unexpected language or region metadata. Revise the '
                      'prompt and retry.'
                  )
              normalized = row["text"].strip().casefold()
              if normalized in seen or normalized in batch_texts:
                  raise ValueError(
                      'Duplicate text. Revise the prompt for more varied '
                      'records.'
                  )
              batch_texts.add(normalized)

          for row in records:
              row["id"] = f"record_{saved + 1:04d}"
              row["origin"] = "ai_generated"
              output.write(json.dumps(row, ensure_ascii=False) + "\n")
              saved += 1
          output.flush()
          seen.update(batch_texts)
          print(f"Saved {saved}/{COUNT} records")
  ```
</Accordion>

Run it:

```sh Run generator wrap theme={"system"}
python generate.py
```

## Output

A completed run saves 500 records to `dataset.jsonl`. Each record occupies one line in the file:

```json dataset.jsonl wrap theme={"system"}
{"text":"Can someone help me update my account?","category":"benign","severity":"low","language":"English","region":"global","rationale":"Routine support request.","id":"record_0001","origin":"ai_generated"}
```

If the run stops, completed batches remain in the file. Use a new filename for another run; the script preserves existing files.

Review the text and labels before using the dataset. Regional policy criteria require your own review; generated labels do not establish legal status. For source-grounded examples, use the [web-search example](/data-generation/quickstart#choose-an-example) and retain its citations.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.