# API Docs

> Interactive reference for the Mirelo AI REST API. Generate sound effects, extend audio, and inpaint segments with a few lines of code.

HTML explorer (canonical for search): https://mirelo.ai/api-docs

Machine-readable OpenAPI v3 spec: https://api.mirelo.ai/v3/openapi.json

This markdown page is the stable agent-facing reference. Endpoint request/response details live in the OpenAPI document; the HTML explorer at `/api-docs` is interactive and not a substitute for this text.

## Overview

Use the Mirelo API to generate sound effects from text or video, extend and edit audio, and transcribe audio into MIDI. This reference describes the endpoints, request parameters, and responses for each workflow.

### Base URL

```text
https://api.mirelo.ai
```

### Start here

Read [Developer Docs](https://mirelo.ai/developers/docs/guides) for API key setup and integration examples. See [Authentication](https://mirelo.ai/api-docs#concept/authentication) for how to authenticate requests and keep your key secure.

### Explore the API

Generate sound effects: [Text to SFX](https://mirelo.ai/api-docs#text-to-sfx) turns a prompt into audio; [Video to SFX](https://mirelo.ai/api-docs#video-to-sfx) generates audio matched to footage.

Edit audio: [Extend Audio](https://mirelo.ai/api-docs#extend-audio) continues an existing clip; [Inpaint Audio](https://mirelo.ai/api-docs#inpaint-audio) regenerates a selected region.

Transcribe music: [Audio-to-MIDI Pro](https://mirelo.ai/api-docs#audio-to-midi) converts audio into MIDI and supports instrument detection.

### Machine-readable references

Use the [API reference (Markdown)](https://mirelo.ai/api-docs.md) for a text version of the documentation, or the [OpenAPI v3](https://api.mirelo.ai/v3/openapi.json) specification for endpoint schemas.

## Authentication

Every request must carry your API key as a Bearer token in the Authorization header. A key is tied to your account and can call every endpoint.

### Getting a key

Go to [Studio → Settings → API Keys](https://mirelo.ai/studio/api-keys). Click "Create key", give it a name, and copy it right away — it is only shown once.

### Using your key

Pass the key as a Bearer token on every request:

`POST https://api.mirelo.ai/v3/text-to-sfx/generations?wait=25`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 8000,
  "input": {
    "prompt": "Thunder crack"
  }
}
```

> **Warning:** Keep keys out of client-side code and version control. Rotate a key right away if it leaks.

### Check your account

GET /v3/me confirms the key works and returns your balance:

`GET https://api.mirelo.ai/v3/me`

```json
{
  "account_type": "user",
  "id": "usr_abc123",
  "email": "you@example.com",
  "credits_available": 840,
  "spend_capacity": 840,
  "recovery_action": null,
  "recovery_url": null,
  "billing_mode": "metered",
  "provisioning_state": "ready",
  "provisioning_deadline": null
}
```

credits_available is your ledger balance. spend_capacity is what the next request may spend: the balance minus credits reserved by your running jobs, plus any overage you have enabled. recovery_action and recovery_url tell you the one step to take when the account cannot pay for anything.

### Organization keys

An organization key is unmetered: usage is logged, not charged to a credit balance, so credits_available and spend_capacity are null. Null here does not mean zero. Check billing_mode before you compare a quote against spend_capacity. email is null too: the organization is the payer, so there is no one user to name.

```json
{
  "account_type": "organization",
  "id": "org_example",
  "email": null,
  "credits_available": null,
  "spend_capacity": null,
  "recovery_action": null,
  "recovery_url": null,
  "billing_mode": "unmetered",
  "provisioning_state": "ready",
  "provisioning_deadline": null
}
```

## Building a product

You can embed Mirelo in your own application so your users generate sound or transcribe audio without leaving your product. This page states the operational facts. The legal rules are in the pages linked below — those pages win if anything here is shorter.

### Plans and commercial use

Any plan can create an API key and call the API. The [Terms](https://mirelo.ai/terms) still apply: the Free plan is for non-commercial use. A paid subscription is what lets you ship a customer-facing product. Compare plans on [Pricing](https://mirelo.ai/pricing).

### What you may build

The [API Terms](https://mirelo.ai/api-terms) let you embed the Services in your product. Do not resell, redistribute, or offer the API as a standalone service. You own Output to the extent intellectual property arises; you must have the rights to the Input you send, including other people's music and voices.

### Attribution

If your users start or choose generations themselves, show a visible "Powered by Mirelo" credit. A fully automatic pipeline — no user initiating or influencing generation — does not need that badge. Free Studio Output has its own attribution rule in the Terms.

### Acceptable use

Your End-Users must follow the [Acceptable Use Policy](https://mirelo.ai/acceptable-use-policy) and the Terms. You are responsible for their use and indemnify Mirelo if they breach those rules. We do not pre-approve applications. We can review content and suspend or terminate access when we have reason to believe the Terms or AUP were broken.

### Training

Paid users can opt out of the training license in [Studio → Settings → Profile](https://mirelo.ai/studio/settings/profile). Free users cannot opt out: the training license is a condition of the free plan. Opting out stops us adding your Input and Output to our training data. Content already added stays there until you delete your account, and uses already underway are not unwound. A member's own setting does not cover organization API keys; email [legal@mirelo.ai](mailto:legal@mirelo.ai) about training on content generated with them. See the [Privacy Policy](https://mirelo.ai/privacy#processing-online-offer-improvement) for your right to object.

### Models and limits

The API and Audio-to-MIDI Pro in Studio use the same Audio-to-MIDI model. See the [Audio-to-MIDI reference](https://mirelo.ai/api-docs#audio-to-midi) for input limits and request options. For sound effects, consult the [API reference](https://mirelo.ai/api-docs#text-to-sfx) and GET /v3/models for available models, supported operations, and duration limits.

### Result availability

An uploaded asset can be used for 24 hours after its upload; after that, upload it again or pass a URL. Expired assets are then deleted from storage automatically. There is no delete route. Async jobs expire after 24 hours. After expiry, results and fresh download links are no longer available through the API. Download-link lifetimes before then vary by endpoint and API version, so follow the response schema and save every result promptly. Generated audio and its record are deleted 7 days after generation, or sooner for some endpoints, and every download link stops working then. If our training-data export is behind, deletion can wait for it for up to 7 more days. API results are not long-term storage. We do not publish an uptime SLA on self-serve. Ask about an enterprise agreement if you need one.

### How we handle misuse

We are not obliged to monitor every request. If you believe stored content is illegal under EU or member-state law, email feedback@mirelo.ai with the location and why. We can remove content and restrict accounts. Circumventing limits, sharing keys, or missing required attribution are grounds to suspend API access.

> **Info:** Treat the API as a generator, not a library. Copy results into your own storage as soon as a job finishes.

## Models

Every create and preflight names a public model release in the body. There is no default or floating latest alias. GET /v3/models lists available releases and their supported operations.

### API version and model release

| Identity | What it selects or records |
| --- | --- |
| Mirelo-Version | The shape and meaning of the HTTP/JSON contract |
| model | A public model release, saved with the result |

```json
{
  "model": "sfx-1.6"
}
```

A job’s model identifies the model release used for that job. Retain it with the output.

### List the models

`GET https://api.mirelo.ai/v3/models`

```json
{
  "data": [
    {
      "id": "sfx-1.6",
      "status": "general_availability",
      "audio": {
        "sample_rate": 44100,
        "channels": 1
      },
      "max_prompt_chars": 5000,
      "max_negative_prompt_chars": 0,
      "stems": [
        "mix",
        "speech",
        "sfx"
      ],
      "controls": [
        "loop",
        "preserve_speech",
        "multi_stem"
      ],
      "formats": [
        "wav",
        "flac",
        "mp3_320"
      ],
      "operations": {
        "text-to-sfx": {
          "duration_ms": {
            "min": 100,
            "max": 60000,
            "min_with_loop": 1000,
            "max_with_loop": 600000
          },
          "num_variants": {
            "min": 1,
            "max": 4
          },
          "controls": [
            "loop"
          ]
        },
        "video-to-sfx": {
          "duration_ms": {
            "min": 1000,
            "max": 600000,
            "max_with_preserve_speech": 60000
          },
          "num_variants": {
            "min": 1,
            "max": 4
          },
          "controls": [
            "preserve_speech",
            "multi_stem"
          ]
        },
        "extend": {
          "append_duration_ms": {
            "min": 1000,
            "min_with_loop": 1000,
            "max": 57000
          },
          "num_variants": {
            "min": 1,
            "max": 4
          },
          "controls": [
            "loop"
          ]
        },
        "inpaint": {
          "region_start_ms": {
            "min": 1000
          },
          "region_width_ms": {
            "min": 1000,
            "max": 8000
          },
          "num_variants": {
            "min": 1,
            "max": 4
          },
          "controls": []
        }
      },
      "credits_per_second": 10
    }
  ]
}
```

GET /v3/models/{id} returns one model. Read operations for the family you are calling: duration limits (including the loop and preserve_speech variants), the num_variants ceiling, and which controls it accepts.

> **Info:** Today the only served model is sfx-1.6. The formats, stems and controls listed on a model are the ones this deployment can actually produce.

### Where to read each limit

| To know | Read |
| --- | --- |
| Which model ids exist | GET /v3/models |
| Duration and num_variants limits | operations.<family> on the model |
| Audio formats you can ask for | formats on the model |
| Whether loop, preserve_speech or multi_stem is available | controls on the model and on the family |
| Which output types can be produced | stems on the model |
| Price per second | credits_per_second on the model |
| How long a prompt can be | max_prompt_chars on the model |
| Whether a negative prompt is accepted, and how long | max_negative_prompt_chars on the model |
| Inputs the model may not follow | operations.<family>.experimental |

### Negative prompts

A negative prompt lists sounds you would rather not hear. Send it as input.negative_prompt on text-to-sfx or video-to-sfx, as a short comma-separated list:

```json
{
  "input": {
    "prompt": "Heavy rain on a metal roof",
    "negative_prompt": "music, speech, thunder"
  }
}
```

Not every model takes one. max_negative_prompt_chars on the model is the longest it accepts, and 0 means it takes none. Sending one to a model that publishes 0 is refused with capability_unsupported, and one over the limit is refused with invalid_request rather than cut short. Leave the field out when you have nothing to exclude.

> **Info:** sfx-1.6 publishes max_negative_prompt_chars 0, so it takes no negative prompt.

### Experimental inputs

Some models accept an input but do not always follow it. operations.<family>.experimental lists those inputs by their request path, such as input.prompt or input.negative_prompt. An experimental input usually steers the result, but check what comes back rather than relying on it.

The list is per model and per family, because the same input can be followed closely by one model and loosely by another. When it is absent, the model has no experimental inputs on that family. An experimental input is still validated like any other.

## Assets & Sources

Endpoints that take audio or video accept the file in two ways: as a public URL, or as an asset you uploaded first. Both give the same result.

### Option A — Public URL

Pass the URL and the server fetches the file. It must be reachable with a plain GET; presigned links work.

```json
{
  "input": {
    "video": {
      "type": "url",
      "url": "https://example.com/my-scene.mp4"
    }
  }
}
```

> **Tip:** URL input is the fastest way to start. Use it when your files are already hosted (S3, a CDN) or for prototyping.

### Option B — Upload an asset first

Ask for an upload ticket, POST the bytes to storage, then pass the asset id. Use this for private files, or to avoid re-fetching a large file on every request.

1. Call POST /v3/assets with the file's content_type: a plain audio/* or video/* type, with no charset or codecs.
2. POST a multipart/form-data body to upload_url. Send every entry of fields first, then the file as the last part, named file. Do not send an Authorization header.
3. Once your upload has returned 2xx, pass the id as { "type": "asset", "id": "…" }. Using it earlier fails with asset_not_ready, which you can retry.

Step 1 — create an upload ticket:

`POST https://api.mirelo.ai/v3/assets`

```json
{
  "content_type": "video/mp4"
}
```

```bash
# Response (201)
{
  "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "upload_url": "https://s3.amazonaws.com/...",
  "upload_expires_at": "2026-09-04T15:00:56.500Z",
  "max_bytes": 2516582400,
  "fields": {
    "key": "ast-1",
    "Policy": "...",
    "X-Amz-Algorithm": "AWS4-HMAC-SHA256"
  }
}
```

Step 2 — POST the file to upload_url. Every fields entry goes first, the file last:

```bash
curl -X POST "https://s3.amazonaws.com/..." \
  -F "key=ast-1" \
  -F "Policy=..." \
  -F "X-Amz-Algorithm=AWS4-HMAC-SHA256" \
  -F "file=@my-scene.mp4"
```

Step 3 — use the id wherever a request takes media:

```json
{
  "input": {
    "video": {
      "type": "asset",
      "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
    }
  }
}
```

> **Info:** The ticket tells you max_bytes and upload_expires_at. A larger or empty body is refused with 400, and an edited field with 403. For a file over max_bytes, host it yourself and pass a URL.

### Which to use?

|  | URL | Asset |
| --- | --- | --- |
| Setup | None | Ticket + multipart POST |
| File must be public | Yes | No |
| Reuse across requests | Any time | Within 24 hours of upload |
| Best for | Prototyping, hosted files | Private files, large files, batch jobs |

## Wait or poll

Every create is one POST. You choose whether the server holds the connection while the job runs, or returns at once so you can poll.

### Wait on the create (up to 25 s)

Add ?wait=25 to the create, or send Prefer: wait=25; the query parameter wins if both are present. Values outside 1–25 are rejected, not clamped. You get 200 with the finished job if it completes inside the window, or 202 with the job still queued or running. The body is the same job object either way, so handle both the same way.

`POST https://api.mirelo.ai/v3/text-to-sfx/generations?wait=25`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 8000,
  "input": {
    "prompt": "Crackling campfire"
  }
}
```

```bash
# 200 — finished inside the window
{
  "id": "gen_01abc123def456",
  "object": "text_to_sfx.generation",
  "status": "succeeded",
  "model": "sfx-1.6",
  "seed": 42,
  "created_at": "2026-09-04T14:00:56.500Z",
  "completed_at": "2026-09-04T14:01:00.200Z",
  "expires_at": "2026-09-05T14:00:56.500Z",
  "credits": 80,
  "errors": [],
  "urls": { "self": "https://api.mirelo.ai/v3/text-to-sfx/generations/gen_01abc123def456" },
  "result": {
    "output": {
      "index": 0,
      "type": "mix",
      "label": null,
      "start_ms": null,
      "end_ms": null,
      "category": null,
      "sound_start_ms": null,
      "sound_end_ms": null,
      "credits": 80,
      "variants_requested": 1,
      "status": "succeeded",
      "gain_db": null,
      "generation_id": "sg_01abc123",
      "variants": [
        {
          "index": 0,
          "status": "succeeded",
          "files": {
            "audio": {
              "url": "https://cdn.mirelo.ai/output/abc123.wav",
              "url_expires_at": "2026-09-04T15:01:00.200Z",
              "bytes": 705644,
              "format": "wav",
              "sample_rate": 44100,
              "channels": 1,
              "duration_ms": 8000
            }
          }
        }
      ]
    }
  }
}
```

> **Tip:** Use wait for short clips and interactive tools. If a job may take longer than 25 s, create without wait and poll.

### Create, then poll

A create without wait returns 202 and the job. Poll GET urls.self until status is final. A poll returns 200 even when the job failed, so branch on status, not on the HTTP code. HTTP errors are separate: 401 for a bad key, 404 once an expired job has been cleaned up, 429 at the rate or concurrency limit.

1. POST to the collection. The response has id, status and urls.self, the absolute URL to poll.
2. GET urls.self every 1–2 seconds. While the job runs, the Retry-After header suggests when to poll next.
3. Stop when status is succeeded, partially_succeeded, failed, canceled or expired.
4. Read result on succeeded or partially_succeeded. Check status on each output and each variant — an entry can be present and still have failed.

```bash
# 202 — accepted
{
  "id": "gen_01abc123def456",
  "object": "text_to_sfx.generation",
  "status": "queued",
  "model": "sfx-1.6",
  "created_at": "2026-09-04T14:00:56.500Z",
  "completed_at": null,
  "expires_at": "2026-09-05T14:00:56.500Z",
  "progress": 0,
  "credits": null,
  "errors": [],
  "urls": { "self": "https://api.mirelo.ai/v3/text-to-sfx/generations/gen_01abc123def456" },
  "result": null
}

# 200 — still running
{
  "id": "gen_01abc123def456",
  "object": "text_to_sfx.generation",
  "status": "running",
  "model": "sfx-1.6",
  "created_at": "2026-09-04T14:00:56.500Z",
  "completed_at": null,
  "expires_at": "2026-09-05T14:00:56.500Z",
  "progress": 0.42,
  "estimated_ms": 3500,
  "credits": null,
  "errors": [],
  "urls": { "self": "https://api.mirelo.ai/v3/text-to-sfx/generations/gen_01abc123def456" },
  "result": null
}
```

> **Info:** Jobs expire 24 hours after creation. A job can only be read with the API key that created it. Use progress and estimated_ms to show progress while it runs.

### Which should I use?

|  | Wait on create | Create, then poll |
| --- | --- | --- |
| Setup | ?wait=1–25 | Polling loop |
| Best for | Short clips, quick prototypes | Long audio, batch, serverless |
| HTTP timeout risk | Yes, if your client times out before the window ends | No |
| Finished body | Job with result | Same job, after poll |

## Credits & Preflight

Generation is charged in credits. On sfx-1.6 the price is 10 credits per second of billed audio, times num_variants. Preflight quotes the charge before you commit; the finished job reports what was actually debited in credits. Read credits_per_second on GET /v3/models rather than hardcoding the rate.

### Your balance

GET /v3/me returns credits_available (your ledger balance) and spend_capacity (what the next request may spend, after reservations held by your running jobs and any overage). Compare a preflight quote against spend_capacity.

`GET https://api.mirelo.ai/v3/me`

### What is billed

| Family | Billed audio |
| --- | --- |
| Text to SFX | duration_ms × num_variants |
| Video to SFX | duration_ms × num_variants |
| Preserve speech | The same, plus 8 credits per second of source audio once when speech is detected |
| Separated stems | The plan's quote, scaled by num_variants |
| Extend | append_duration_ms × num_variants — the source is never billed |
| Inpaint | region width × num_variants — the rest of the file is never billed |
| Audio-to-MIDI Pro | 2.5 credits per second of input audio (rounded to the nearest credit, minimum 1) |

### Preflight: POST the same body as create

Every collection has a POST …/preflight route. Send the create body and you get the credit cost and the estimated processing time. Nothing is charged and no job starts.

`POST https://api.mirelo.ai/v3/text-to-sfx/generations/preflight`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 8000,
  "num_variants": 1,
  "input": {
    "prompt": "Heavy rain on a metal roof"
  }
}
```

```bash
# Response
{
  "credits": 80,
  "estimated_ms": 4200,
  "credit_recovery": {
    "credits_required": 80,
    "credits_available": 840,
    "credit_shortfall": 0,
    "recovery_action": null,
    "recovery_url": null,
    "provisioning_state": "ready",
    "provisioning_deadline": null
  }
}
```

> **Tip:** On a metered key, check credit_recovery before you create: provisioning_state must be "ready", recovery_action must be null, and credit_shortfall must be 0. Preflight does not reserve credits, so the create can still answer 402 if the balance changes in between.

### Unmetered keys

On an unmetered key (billing_mode "unmetered" on GET /v3/me), credits in a quote is the list price for reporting, not a charge. credit_recovery is omitted, and the finished job reports credits: null. Do not compare the quote against spend_capacity, which is null too.

### Partial and failed jobs

Billing is all-or-nothing on the variant count you requested. A partially_succeeded job reports the full charge, and a failed job can too: result_unreadable means audio was generated and billed but could not be read back, and that error is not retryable. See Failed & partial jobs.

### 402 and credit_recovery

A create that cannot be paid answers 402. The error carries the same credit_recovery object preflight would have shown. Use it to decide the next step; do not parse the message.

```json
{
  "error": {
    "code": "insufficient_credits",
    "message": "This request needs 20 credits; 8 are available.",
    "param": null,
    "retryable": false,
    "request_id": "req_01abc123def456",
    "credit_recovery": {
      "credits_required": 20,
      "credits_available": 8,
      "credit_shortfall": 12,
      "recovery_action": "enable_overage",
      "recovery_url": "https://mirelo.ai/studio/settings/billing",
      "provisioning_state": "ready",
      "provisioning_deadline": null
    }
  }
}
```

## How to preserve speech

Keep the dialogue already in the clip. Same video-to-sfx collection: set controls.preserve_speech and the clip's own soundtrack is read. Generated sound is placed around the speech and ducked under it.

1. Send the video on input.video. Add input.audio only when the audio to preserve is a separate file, and keep it aligned with the clip: start_offset_ms selects the same window from both.
2. Set controls.preserve_speech to true.
3. Stay within the lower duration ceiling (max_with_preserve_speech: 60 s on sfx-1.6). Read it from GET /v3/models rather than hardcoding it.
4. Poll the job. When speech is found, result.outputs is mix, sfx, and speech.

### Request

`POST https://api.mirelo.ai/v3/video-to-sfx/generations?wait=25`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 10000,
  "input": {
    "video": {
      "type": "url",
      "url": "https://example.com/scene.mp4"
    },
    "prompt": "Classroom chatter under a pencil tap"
  },
  "controls": {
    "preserve_speech": true
  }
}
```

> **Info:** Preflight quotes the upper-bound price. The settled charge drops to the base video-to-sfx price when no speech is detected.

> **Warning:** A clip with no audio track at all is refused with invalid_video rather than coming back having preserved nothing. Preserve speech cannot be combined with controls.multi_stem — that control already returns a fixed mix / sfx / speech triple. The speech output is your own audio rather than generated audio, and is never persisted for feedback.

### What comes back

Three outputs when speech is found: mix (speech under SFX), sfx (generated only), and speech (isolated dialogue). If nothing was detected, the speech output is omitted.

## How to generate separate stems

Magic mode detects sound events and generates an independent clip for each. It does not split an existing mix into stems. It takes two requests: preflight detects the sounds and quotes a plan, then create runs that plan.

1. POST the create body, with controls.multi_stem set to true and no plan, to the preflight route. The response lists the detected outputs and includes plan_id and plan_expires_at.
2. Check the quote. Each entry in outputs has a label, a category (foreground, background or swoosh), start_ms, end_ms, sound_start_ms, sound_end_ms and its own credits.
3. POST the same body to the collection with plan set to that plan_id. Keep model, input, duration_ms, start_offset_ms and controls.multi_stem unchanged. num_variants may change; the charge scales with it.
4. Poll the job. Each detected sound is its own entry in result.outputs, with the same index, label, category and timing fields as the plan, so you can place it on your timeline.

### Step 1 — preflight with stems

`POST https://api.mirelo.ai/v3/video-to-sfx/generations/preflight`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 10000,
  "input": {
    "video": {
      "type": "url",
      "url": "https://example.com/scene.mp4"
    }
  },
  "controls": {
    "multi_stem": true
  }
}
```

```bash
# Response
{
  "credits": 71,
  "estimated_ms": 18000,
  "plan_id": "plan_4Qb8vT2nHk",
  "plan_expires_at": "2026-09-05T14:00:56.500Z",
  "outputs": [
    {
      "index": 0,
      "type": "sfx",
      "label": "footsteps",
      "start_ms": 400,
      "end_ms": 4400,
      "category": "foreground",
      "sound_start_ms": 1400,
      "sound_end_ms": 3400,
      "credits": 40,
      "variants_requested": 1
    },
    {
      "index": 1,
      "type": "sfx",
      "label": "door close",
      "start_ms": 4100,
      "end_ms": 7200,
      "category": "foreground",
      "sound_start_ms": 5100,
      "sound_end_ms": 6200,
      "credits": 31,
      "variants_requested": 1
    }
  ],
  "credit_recovery": {
    "credits_required": 71,
    "credits_available": 840,
    "credit_shortfall": 0,
    "recovery_action": null,
    "recovery_url": null,
    "provisioning_state": "ready",
    "provisioning_deadline": null
  }
}
```

### Step 2 — create with the plan

`POST https://api.mirelo.ai/v3/video-to-sfx/generations?wait=25`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 10000,
  "input": {
    "video": {
      "type": "url",
      "url": "https://example.com/scene.mp4"
    }
  },
  "controls": {
    "multi_stem": true
  },
  "plan": "plan_4Qb8vT2nHk"
}
```

> **Info:** start_ms and end_ms are the generated window, which reaches up to a second past the sound on each side (less near the start or end of the video) so you can crossfade it into its neighbours. sound_start_ms and sound_end_ms mark the sound itself: trim the clip to that span. Treat category as an open list and handle a value you do not recognise.

> **Info:** Rules: omit input.prompt and seed, which this mode cannot honor. Do not combine with preserve_speech. Stems always come back as wav (another output.format is refused), start_offset_ms must stay 0, and a create that asks for stems without a plan is refused. A plan is valid for 24 hours and only for the body it was quoted for. If the body differs, the create fails with invalid_request on plan; an unknown or expired plan fails with not_found.

> **Warning:** Preflight runs detection but does not reserve credits, so the create can still answer 402. If outputs is empty, nothing was detected and there is nothing to generate.

## Idempotency

Creates accept an optional Idempotency-Key header. Send the same key with the same body again and you get the original job back instead of a second generation. The same key with a different body is refused with 409 idempotency_conflict.

### When to send one

Send a key on any create you might retry: a client timeout, a 5xx, or a connection that dropped while you were waiting. Preflight, poll, models, me and assets do not take one.

> **Info:** The key is yours to choose; a UUID is enough. A key is scoped to your account and to the collection you sent it to, and is remembered for 24 hours.

### Retry after a lost response

Save the key and the exact body before you send the create. If the response is lost, send the same request again. You get the original job — 200 if it has finished, 202 if it is still running — and the Idempotent-Replayed: true header tells you it was a replay. Poll urls.self as usual.

`POST https://api.mirelo.ai/v3/text-to-sfx/generations?wait=25`

Header: `Idempotency-Key: 8f3c1e2a-5b7d-4c9e-9a1f-2d3e4f5a6b7c`

```json
{
  "model": "sfx-1.6",
  "duration_ms": 8000,
  "input": {
    "prompt": "Heavy rain on a metal roof"
  }
}
```

> **Warning:** Do not mint a new key just because a response was lost: that starts a second paid job. A 409 means the body changed, so find the original body and resend that. A job that already failed is not repaired by replaying its key. Starting a new generation is a new decision and may be charged again.

## Failed & partial jobs

A 200 from a poll means the API returned the job, not that the job succeeded. Branch on status. Request problems come back in error; problems inside an accepted job come back in errors on the job.

### Partial success

This job asked for two variants. One is downloadable; the other was generated but could not be read back. Both variants are listed, so do not compare variants.length to variants_requested — read status on the output and on each variant. The full charge for two variants remains.

```json
{
  "id": "gen_partial_example",
  "object": "text_to_sfx.generation",
  "status": "partially_succeeded",
  "model": "sfx-1.6",
  "created_at": "2026-09-04T14:00:56.500Z",
  "completed_at": "2026-09-04T14:01:00.200Z",
  "expires_at": "2026-09-05T14:00:56.500Z",
  "credits": 160,
  "errors": [{
    "output_index": 0,
    "variant_index": 1,
    "code": "result_unreadable",
    "message": "The generated file could not be read back.",
    "retryable": false
  }],
  "urls": { "self": "https://api.mirelo.ai/v3/text-to-sfx/generations/gen_partial_example" },
  "result": {
    "output": {
      "index": 0,
      "type": "mix",
      "label": null,
      "start_ms": null,
      "end_ms": null,
      "category": null,
      "sound_start_ms": null,
      "sound_end_ms": null,
      "credits": 160,
      "variants_requested": 2,
      "status": "partially_succeeded",
      "gain_db": null,
      "generation_id": "sg_partial_example",
      "variants": [
        {
          "index": 0,
          "status": "succeeded",
          "files": { "audio": {
            "url": "https://cdn.mirelo.ai/output/partial.wav",
            "url_expires_at": "2026-09-04T15:01:00.200Z",
            "bytes": 705644,
            "format": "wav",
            "sample_rate": 44100,
            "channels": 1,
            "duration_ms": 8000
          } }
        },
        { "index": 1, "status": "failed", "files": null }
      ]
    }
  }
}
```

### A failed job can still be charged

Here audio was generated and debited, but the file could not be read back. result_unreadable is not retryable: a retry would generate and charge a second copy. Keep the job id and errors for support. When nothing was debited, credits is 0. On an organization key credits is always null.

```json
{
  "id": "gen_failed_example",
  "object": "text_to_sfx.generation",
  "status": "failed",
  "model": "sfx-1.6",
  "created_at": "2026-09-04T14:00:56.500Z",
  "completed_at": "2026-09-04T14:01:00.200Z",
  "expires_at": "2026-09-05T14:00:56.500Z",
  "credits": 80,
  "errors": [{
    "output_index": null,
    "variant_index": null,
    "code": "result_unreadable",
    "message": "Audio was generated and billed but could not be read back.",
    "retryable": false
  }],
  "urls": { "self": "https://api.mirelo.ai/v3/text-to-sfx/generations/gen_failed_example" },
  "result": null
}
```

### What to do next

| Outcome | Action |
| --- | --- |
| partially_succeeded | Download the variants whose status is succeeded. Report the others using output_index and variant_index from errors. |
| failed, retryable false | Stop. Fix the input or contact support with the job id. Replaying the same idempotency key returns this same failed job. |
| failed, retryable true | You may send the request again as a new create with a new key. It is billed as new work. |
| canceled or expired | Stop polling. Anything you already downloaded is still yours; regenerating is a new job. |

## Downloads & expiry

Download links are temporary. Save the audio in your own storage as soon as a job finishes.

### Get a fresh download link

If a link has expired but the job has not, GET urls.self again with the same API key. The response carries fresh links for the same outputs and variants. A GET is a read: it starts no work and costs no credits.

`GET https://api.mirelo.ai/v3/text-to-sfx/generations/gen_01abc123def456`

### Lifetimes

| Field | What it covers | When it ends |
| --- | --- | --- |
| Job — expires_at | The job and its results. 24 hours from creation. | The results are gone, and GET may return 404 once the job is cleaned up. |
| Generated audio | The stored audio behind every download link. 7 days from generation, or sooner for some endpoints; up to 7 more days while the training-data export catches up. | The audio and its record are deleted; download links return 404. |
| File — url_expires_at | One download URL. Read the timestamp rather than assuming a fixed lifetime. | GET the job again for a new link, any time before the job expires. |
| Plan — plan_expires_at | A separated-stems plan. 24 hours from preflight. | Preflight again for a new plan. |
| Upload — upload_expires_at | The window in which an asset upload may start. Currently one hour. | Mint a new ticket with POST /v3/assets and upload again. |

> **Info:** A download link may redirect to storage. Follow the redirect, and do not send your API key with it — the link authorizes itself.

Audio-to-MIDI download URLs are temporary too. Save every result promptly; after the job expires, you cannot poll it for fresh links. See [Building a product](https://mirelo.ai/api-docs#concept/building-on-the-api) for integration guidance.

## Errors & Limits

Every error uses the same envelope. Treat code as an open list: a client should tolerate a value it does not know.

```json
{
  "error": {
    "code": "invalid_request",
    "message": "num_variants is above the model's ceiling.",
    "param": "num_variants",
    "retryable": false,
    "request_id": "req_01abc123def456"
  }
}
```

| Field | Meaning |
| --- | --- |
| error.code | Stable machine-readable code |
| error.message | For humans; do not parse it |
| error.param | The JSON body path that failed, or null |
| error.retryable | Whether the same request may succeed later |
| error.request_id | Quote this when you contact support |

### Concurrency and rate limits

By default, up to 5 of your jobs run at the same time. The ceiling is set per key, so read Mirelo-Concurrency-Limit rather than assuming the default. Create and poll responses carry Mirelo-Concurrency-Current and Mirelo-Concurrency-Limit. At the ceiling a create answers 429 with those headers and a Retry-After. The RateLimit-* headers describe the request-rate allowance; for generation, concurrency is the limit you reach first.

### Version pin

Send Mirelo-Version to pin the API behaviour you tested against. The current version date is in the OpenAPI document. Without the header you get the latest.

### Whole milliseconds

Every duration field is a whole number of milliseconds. Fractional values are rejected.

## Hosted MCP

Mirelo runs an official, remotely hosted Model Context Protocol (MCP) server so AI agents — Cursor, Claude, and other MCP-capable clients — can generate audio directly from chat. Connect once, sign in with your existing Mirelo account, and the agent can check your credits, generate SFX, and poll long-running jobs.

### How it works

Connecting uses Sign in with Mirelo (Clerk OAuth) — the same account as Studio. There is no per-user sk- API key to create, paste, or rotate for this flow, and no separate SDK on the generate path. Access is available to signed-in Studio accounts (free and paid); every tool call is billed against your normal Studio credits, exactly like generating from the Studio UI or the public API.

### Connect

Copy this URL into your agent's remote MCP / connector settings:

```bash
https://mcp.mirelo.ai/mcp
```

#### Cursor

```json
{
  "mcpServers": {
    "mirelo": {
      "url": "https://mcp.mirelo.ai/mcp"
    }
  }
}
```

1. Add the JSON above to Cursor's MCP settings (or add a remote server pointing at the URL).
2. Click Connect, then complete Sign in with Mirelo.
3. Tools appear in the agent once OAuth succeeds.

#### Claude and other clients

Use the same https://mcp.mirelo.ai/mcp URL with your client's remote MCP / connector flow. If a client only supports stdio, wrap it with mcp-remote pointing at that URL (streamable HTTP; no extra flags needed over HTTPS).

### Tools

| Tool | Mode | Notes |
| --- | --- | --- |
| get_account | — | Credits and account email |
| preflight | — | Credit cost + ETA for any SFX generate/edit action via endpoint_key |
| create_upload | — | Preferred way to send a local file. Returns a pre-signed URL — the agent PUTs the file directly, so bytes never pass through the conversation |
| inspect_asset | — | Read a create_upload asset's real duration, content type, and size before generating |
| create_asset | — | Inline upload for very small files only (bytes ride in the tool call). Prefer create_upload for anything of real size |
| text_to_sfx | async | Returns a job_id |
| video_to_sfx | async | Returns a job_id; video accepts url or asset_id |
| extend_audio | async | Returns a job_id; audio accepts url or asset_id |
| extend_audio_with_video | async | Returns a job_id |
| inpaint_audio | async | Optional video guide; returns a job_id |
| audio_to_midi | async | Transcribe a recording into MIDI and MusicXML, with optional engraved PDFs; returns a job_id |
| get_transcription_notes | — | Read a finished transcription's notes in pages. Only needed for note-level work; the job result is a compact summary |
| get_job | async | Client-driven poll for a job_id (preferred when the agent can loop) |
| wait_job | async | Block until a job_id finishes |

### Async jobs

All generate/edit tools are async-only. Call text_to_sfx, video_to_sfx, extend_audio, extend_audio_with_video, inpaint_audio, or audio_to_midi to get a job_id, then get_job to poll (or wait_job to block until it finishes). Prefer get_job when the agent can poll itself.

> **Info:** Current scope: music, multi-stem, Roblox tools, and Studio projects/timeline editing are deferred and not exposed over MCP.

## Migrate from V2

This page maps V2 onto V3. Everywhere else, these docs describe V3 on its own terms. Audio-to-MIDI has no V3 collection yet: those paths stay /v2/audio-to-midi/… and are unchanged.

### Paths

| V2 | V3 |
| --- | --- |
| /v2/text-to-sfx/v1.6/sync or /jobs | POST /v3/text-to-sfx/generations |
| /v2/video-to-sfx/v1.6/… | POST /v3/video-to-sfx/generations |
| /v2/video-to-sfx-preserve-speech/v1.6/… | Same video-to-sfx collection, with controls.preserve_speech |
| /v2/extend-audio/v1.6/… and /with_video/… | POST /v3/extend/edits (video is the optional input.video) |
| /v2/inpaint-audio/v1.6/… and /with_video/… | POST /v3/inpaint/edits (video is the optional input.video) |
| GET …/preflight?duration_ms=&num_samples= | POST …/preflight with the same body as create |
| /v2/video-to-sfx-stems/v1.6/analyze and /jobs | Video-to-SFX preflight and generations, with controls.multi_stem: true and plan on create |
| GET /v2/me | GET /v3/me |
| POST /v2/assets, then PUT to upload_url | POST /v3/assets, then a multipart POST with fields + file |
| openapi.yaml | GET /v3/openapi.json |

### Request shape

| V2 | V3 |
| --- | --- |
| Model in the URL (/v1.6/) | Required body field model (sfx-1.6) |
| Flat prompt, duration_ms, loop, output_format | Nested input / controls / output |
| num_samples | num_variants |
| output_format: wav \| mp3 \| aac \| flac | output.format: wav, flac, mp3_320 |
| { type: "url", video_url / audio_url } or { type: "asset", asset_id } | { type: "url", url } or { type: "asset", id } |
| seed -1 or null for random | Omit seed; -1 is rejected |
| Sync POST …/sync | POST the collection with ?wait=25 (or Prefer: wait=25) |
| Async POST …/jobs, then GET …/jobs/{job_id} | POST the collection, then GET …/{id}; a poll is always 200 |

### Response shape

| V2 | V3 |
| --- | --- |
| result_urls[] | Job envelope: id, status, result.output(s).variants[].files.audio |
| job_id / job_url / errored | id / urls.self / failed |
| estimated_completion_at, request echo | Gone — use estimated_ms, progress, and your own records |
| Partial failure via shorter arrays | partially_succeeded; check status on each output and variant |
| Flat error or status text | { error: { code, message, param, retryable, request_id } } |
| 402 without structured recovery | 402 includes credit_recovery |

### Features that fold into one collection

Preserve speech is no longer its own path: set controls.preserve_speech on video-to-sfx, and send input.audio only when the audio to preserve is not the clip's own. Extend-with-video and inpaint-with-video are the same edits collections with an optional input.video. On extend, loop and video cannot be combined, prepend_duration_ms is refused, and start_offset_ms needs a video.

Multi-stem generation uses the same planner as V2: preflight with controls.multi_stem: true, then create with the returned plan_id as plan. Obtain a new V3 plan when migrating from V2; V2 plan ids cannot be used on V3. See how to generate separate stems.

### Headers

Send Idempotency-Key on creates you might retry. Pin behaviour with Mirelo-Version. Read Mirelo-Concurrency-* (default ceiling 5) and RateLimit-*.

### Machine-readable spec

The V3 OpenAPI document is GET /v3/openapi.json. It is public and not behind a flag. Use it to generate clients; these pages stay a hand-written reference so examples and limits cannot drift from the constants the server enforces. Audio-to-MIDI is included on its existing /v2/audio-to-midi/… paths until it has a V3 collection.

## Billing & retries

### Preflight prices your duration estimate

GET /v2/audio-to-midi/v1.0/preflight applies max(1, round(seconds × 2.5)) to the duration_ms you submit, converted to seconds. It does not download or inspect the audio and charges nothing.

`GET https://api.mirelo.ai/v2/audio-to-midi/v1.0/preflight?duration_ms=300000`

### Final billing uses the measured audio

The sync and async create routes download the media, probe its real duration, and apply the same formula to that. A preflight for 300 seconds quotes 750 credits; media that measures 308.059 seconds bills 770. Keep the submitted duration accurate and treat the quote as an estimate.

### Retry synchronous requests safely

Send Idempotency-Key (up to 128 characters) on every sync request you might retry: one stable key per logical transcription, the same body every time, reused only when the first response was ambiguous, such as after a client timeout. A completed request's key is remembered for 24 to 48 hours after the original request completes, depending on when the daily sweep runs. A 409 retry does not extend this window.

```bash
operation_key="order-123"

transcribe() {
  curl --fail-with-body --max-time 600 \
    https://api.mirelo.ai/v2/audio-to-midi/v1.0/sync \
    -H "Authorization: Bearer $MIRELO_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: $operation_key" \
    --data '{"audio":{"type":"url","audio_url":"https://example.com/song.mp3"}}'
}

transcribe
# If the response was ambiguous, retry the identical request with the same key.
transcribe
```

| Key reuse | Result |
| --- | --- |
| Same body; first request still running | 409 idempotency_conflict |
| Same body; first request completed and still remembered | 409 idempotency_conflict |
| Different body | 409 idempotency_conflict |
| First attempt failed | The request runs again under the same key |
| Any reuse after the key is forgotten (24 to 48 hours after the original request completes) | A new request runs and is billed |

> **Warning:** A 409 means the retry did not start a second billable transcription, but it does not replay the original response or extend the key's retention window, so keep any success you receive. Sync accepts audio up to 12 minutes; longer recordings return 422 sync_too_long before transcription or billing. Use POST /v2/audio-to-midi/v1.0/jobs instead. A longer client timeout does not extend the server runtime.

### Prefer async for long or high-volume work

Use POST /v2/audio-to-midi/v1.0/jobs for long files, batches, serverless functions, and pipelines that need a retrievable job_id instead of one long-lived connection, then poll job_url until the job succeeds or errors. Idempotency-Key is not supported on the async create route: resubmitting POST .../jobs can create and bill a separate job.

## Instrument detection

The instruments field is exhaustive: an instrument left out cannot appear in the transcription, and one that is not playing pulls notes onto the wrong part. POST /v2/audio-to-midi/v1.0/instruments listens to sampled windows of your audio and suggests a list to start from.

### Detect, review, then transcribe

1. Upload the audio once as an asset, the same kind of asset_id an Audio-to-MIDI request takes.
2. Call POST /v2/audio-to-midi/v1.0/instruments with that asset. A URL source works too; the response returns the asset_id of the stored copy.
3. Review the suggestions and correct the list.
4. Transcribe with the same asset, your corrected instruments, and the instrument_detection_id.

`POST https://api.mirelo.ai/v2/audio-to-midi/v1.0/instruments`

```json
{
  "audio": {
    "type": "asset",
    "asset_id": "3f1f4c1e-8a4b-4b8e-9f9e-2b6c1d0a7e55"
  },
  "max_credits": 0
}
```

### Suggestions need review

recommended_instruments is the detector's best list, ready to pass as instruments. instruments lists everything it admitted, with agreement: the largest share of one window's independent listens that named the instrument. Agreement is a rough signal of how often a label is right, not a probability. An instrument the detector missed carries no score at all.

> **Warning:** Review every suggestion, not only the ones with low agreement, and add anything that is missing. There are no timestamps: the transcription result says exactly when each instrument plays.

### Pricing

| Situation | Charge |
| --- | --- |
| A transcription that succeeds passes the instrument_detection_id | Free |
| Fewer than 10 unused detections on your account in the last 24 hours | Free |
| Past that allowance | Up to 50 credits, never more than transcribing the same audio; they come off the transcription that uses the detection within 30 days, in the same billing period |
| The same audio again, under the same detector revision | Free; returns the cached answer |

A detection is charged only when it succeeds. The allowance counts per account, across all of your API keys. Each response reports credits_charged and free_detections_remaining.

> **Tip:** Send max_credits: 0 to run a detection only when it is free. Past the allowance it fails with 402 max_credits_exceeded and charges nothing, so an app can detect as soon as a user opens it and offer the paid detection only when they ask.
