# Audio (https://hypit.ai/api-reference/audio/)

> Speech synthesis and transcription are synchronous; music generation is a job.

Audio endpoints have different response shapes:

| Route                                      | Shape                                                   |
| ------------------------------------------ | ------------------------------------------------------- |
| `POST /v1/audio/speech`                    | **synchronous** — returns audio bytes, or JSON          |
| `POST /v1/audio/transcriptions`            | **synchronous** — returns text or JSON                  |
| `GET /v1/audio/transcriptions/:request_id` | **read-only** — retrieves a saved transcription         |
| `POST /v1/audio/music`                     | **asynchronous** — returns a [job](/api-reference/jobs) |

<Callout type="info" title="Check availability first">
  Audio models are enabled per deployment. Find the ones your key can reach before you write against
  a name:

  ```bash
  curl -s https://hypit.ai/v1/models \
    -H "Authorization: Bearer $HYPIT_API_KEY" |
    jq -r '.data[] | select(.endpoints[] | test("audio|transcriptions")) | "\(.id)\t\(.endpoints)"'
  ```
</Callout>

## Speech [#speech]

```bash
curl https://hypit.ai/v1/audio/speech \
  -H "Authorization: Bearer $HYPIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "'"$MODEL"'", "input": "The boat went over the weir.", "voice": "alloy"}' \
  --output speech.mp3
```

`client.audio.speech.*` in the OpenAI SDKs works against this route unchanged.

### Request body [#request-body]

| Field                                | Type             | Notes                                                                    |
| ------------------------------------ | ---------------- | ------------------------------------------------------------------------ |
| `model`                              | string           | **required**                                                             |
| `input`                              | string           | the text. `text` is accepted as a synonym                                |
| `prompt`                             | string           | free-form direction, on models that take one                             |
| `lyrics`                             | string           |                                                                          |
| `voice`                              | string           | a named voice                                                            |
| `reference_id`                       | string           | a cloned-voice id                                                        |
| `voice_description`                  | string           | describe a voice instead of naming one                                   |
| `response_format`                    | string           | `mp3`, `wav`, `pcm`, `opus`, `flac`, `aac`, `ogg`. `format` is a synonym |
| `speed`, `volume`, `loudness`        | number           | vendor-defined ranges                                                    |
| `sample_rate`, `bitrate`             | integer          | `mp3_bitrate` is a synonym for `bitrate`                                 |
| `duration_seconds`                   | number           | `seconds` and `duration` are synonyms                                    |
| `n`                                  | integer          | 1–8, default `1`                                                         |
| `seed`                               | integer          |                                                                          |
| `language`                           | string           |                                                                          |
| `instrumental`, `loop`               | boolean          |                                                                          |
| `guidance_scale`, `prompt_influence` | number           |                                                                          |
| `reference_audio`                    | array of strings | `http(s)` or `data:` URLs                                                |
| `auto_generate_text`                 | boolean          |                                                                          |
| `output`                             | string           | `binary` (default), `b64_json`, or `url`                                 |

At least one of `input`/`text`, `prompt`, `lyrics` or `voice_description` is required. Combined
`input` + `lyrics` is capped at **40 000 characters**; the whole JSON body at **12 MiB**. Unmodelled
keys are forwarded to the upstream verbatim.

### MiMo VoiceClone [#mimo-voiceclone]

`mimo-v2.5-tts-voiceclone` requires exactly one MP3 or WAV voice sample. Prefer
`reference_audio`; it accepts either an `http(s)` URL or a complete data URL, and the gateway fetches
and inlines it when the selected provider requires bytes. If you use `voice` directly, send a data URL
(bare base64 from older clients is also accepted):

```json
{
  "model": "mimo-v2.5-tts-voiceclone",
  "input": "This sentence is spoken in the voice from the reference audio.",
  "reference_audio": ["data:audio/wav;base64,UklGRg..."],
  "response_format": "wav"
}
```

MiMo requires `data:audio/mpeg;base64,...` or `data:audio/wav;base64,...`, with at most 10 MiB in
the base64 portion. The gateway checks the file's magic bytes before dispatch, so a spoofed MIME type
or another audio container returns 400 without spending an upstream call.

### Response [#response]

With `output` unset or `binary`, the response is **raw audio bytes** with the sniffed
`Content-Type` (defaulting to `audio/mpeg`), `X-Content-Type-Options: nosniff` and a
`Content-Disposition: inline; filename="job_….mp3"`.

With `output` set to `b64_json` or `url`, the response is JSON:

```json
{
  "object": "audio.speech",
  "created": 1787824589,
  "model": "…",
  "format": "mp3",
  "mime_type": "audio/mpeg",
  "b64_json": "SUQzBA…",
  "seconds": 3.4,
  "usage": { "characters": 28, "audio_seconds": 3.4 }
}
```

`output: "url"` quietly falls back to `b64_json` when no stored URL could be produced, so handle both
keys. A voice-design request answers with `"object": "audio.voice_previews"` and a `previews[]`
array of `{voice_id, name, description, b64_json|url, mime_type, format, seconds}`.

## Transcriptions [#transcriptions]

```bash
curl https://hypit.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $HYPIT_API_KEY" \
  -F model="$MODEL" \
  -F file=@interview.mp3 \
  -F response_format=verbose_json
```

`client.audio.transcriptions.*` in the OpenAI SDKs works against this route unchanged.

Accepts `multipart/form-data` **or** `application/json`. Anything else is
`400 unsupported_content_type`.

| Field                       | Notes                                                                                           |
| --------------------------- | ----------------------------------------------------------------------------------------------- |
| `model`                     | **required**                                                                                    |
| `file`                      | the clip — multipart only                                                                       |
| `url`                       | an `http(s)` URL instead of a file. Exactly one of `file` / `url`                               |
| `response_format`           | `json` (default), `text`, `verbose_json`                                                        |
| `language`                  | ISO code hint                                                                                   |
| `prompt`                    | context hint                                                                                    |
| `temperature`               | number                                                                                          |
| `timestamp_granularities[]` | repeated: `segment` and/or `word`. The bare spelling `timestamp_granularities` is also accepted |

Limits: **100 MiB** per uploaded file, **120 MiB** for the whole multipart body, **4 MiB** for a JSON
body. The declared `Content-Type` of the part is ignored — the bytes are sniffed, and the accepted
containers are WAV, FLAC, Ogg (Vorbis/Opus), MP3, ADTS AAC, M4A, MP4 and WebM. Raw headerless PCM is
rejected.

`response_format: "text"` returns `text/plain` with the bare transcript. Everything else returns
JSON:

```json
{
  "task": "transcribe",
  "language": "en",
  "duration": 92.4,
  "text": "…",
  "segments": [
    { "id": 0, "start": 0.0, "end": 3.2, "text": "…", "no_speech_prob": 0.01 }
  ],
  "words": [{ "word": "the", "start": 0.10, "end": 0.22 }],
  "usage": { "audio_seconds": 92.4 }
}
```

`task`, `language`, `duration`, `segments` and `words` only appear for `verbose_json`. The `usage`
block is ours and is present on the plain `json` shape too — unlike OpenAI, which omits it.

### Retrieve a saved transcription [#retrieve-a-saved-transcription]

Save the server-generated `X-Request-Id` and `Location` response headers from the transcription
POST. Once its durable reservation exists, `Location` points to this read-only endpoint, including
when the POST later reports an uncertain upstream result:

```bash
curl https://hypit.ai/v1/audio/transcriptions/req_YOUR_SERVER_REQUEST_ID \
  -H "Authorization: Bearer $HYPIT_API_KEY"
```

Use a currently valid API key belonging to the original account. OAuth tokens need `user:jobs`;
`user:inference:audio` alone does not permit this read. Historical results remain readable if the model
is later removed or the key's groups change. Another account receives the same 404 as an unknown ID.
Reading never submits another upstream request or charges again.

```json
{
  "id": "req_YOUR_SERVER_REQUEST_ID",
  "object": "audio.transcription",
  "model": "your-model",
  "status": "completed",
  "settlement_state": "settled",
  "created_at": 1788900000,
  "completed_at": 1788900010,
  "expires_at": 1791492010,
  "result": {
    "text": "Hello world.",
    "language": "en",
    "duration": 4.5,
    "usage": { "audio_seconds": 4.5 }
  }
}
```

The saved `result` is JSON regardless of the POST's `response_format`; `segments` and `words` are
included when supplied by the provider. Provider raw responses and input URLs are not exposed.

| HTTP | `status`                | Meaning                                                                       |
| ---- | ----------------------- | ----------------------------------------------------------------------------- |
| 202  | `pending`, `processing` | No saved result yet; wait at least `Retry-After` seconds before reading again |
| 200  | `completed`             | The saved transcript is in `result`                                           |
| 200  | `failed`                | The upstream operation definitively failed                                    |
| 200  | `unavailable`           | Processing ended without a reliable saved transcript                          |
| 410  | `expired`               | The transcript's retention period ended                                       |

`status` describes the transcript; `settlement_state` describes billing independently. A recovered
transcript can be `completed` while historical billing remains `unknown`. Normal successful POSTs
save their results before sending the body. Accepted Replicate transcription tasks can also be
polled after a restart using their saved upstream task ID; this does not promise background result
recovery for every provider or every failure.

Results use the deployment's storage retention period: 30 days by default, measured from first
result persistence; a negative retention setting keeps them indefinitely. Expiry removes the
transcript while retaining financial records. Responses use `Cache-Control: private, no-store`.
The saved public result must fit the same 32 MiB limit as buffered upstream responses; oversized
results are rejected, never silently truncated. A confirmed oversized, undeliverable result returns
502 and releases the customer reservation; it is exposed as `unavailable`, with zero customer charge.
An ordinary client disconnect or a missing GET body is not evidence for a refund.

This route requires the **server-generated** request ID. If the connection drops before those
headers arrive, the client cannot discover that ID through this endpoint. A client-supplied
`X-Request-Id` is not an idempotency key. Repeating the POST creates another request and can charge
again; GET recovery does not add POST idempotency.

## Music [#music]

`POST /v1/audio/music` is asynchronous. It accepts exactly the same body as `/v1/audio/speech` —
same fields, same synonyms, same validation, same 12 MiB cap — and answers `202 Accepted` with a job
whose `kind` is `audio`.

```bash
curl https://hypit.ai/v1/audio/music \
  -H "Authorization: Bearer $HYPIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "'"$MODEL"'",
    "prompt": "slow lo-fi piano, rain outside a window",
    "lyrics": "",
    "instrumental": true,
    "duration_seconds": 60
  }'
```

The fields that matter here are `prompt`, `lyrics`, `instrumental`, `duration_seconds`, `voice`,
`reference_id`, `reference_audio`, `language`, `seed` and `n`. The `output` field is parsed and
validated but has no effect: a job's artifacts always arrive through the assets API.

Then poll and collect exactly as in [Asynchronous jobs](/api-reference/jobs). Two aliases point at
the same handlers:

```bash
curl https://hypit.ai/v1/audio/music/$JOB_ID \
  -H "Authorization: Bearer $HYPIT_API_KEY"

curl -L -o track.mp3 https://hypit.ai/v1/audio/music/$JOB_ID/content \
  -H "Authorization: Bearer $HYPIT_API_KEY"
```

## SDK examples [#sdk-examples]

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["HYPIT_API_KEY"],
    base_url=os.environ.get("HYPIT_BASE_URL", "https://hypit.ai/v1"),
)

with client.audio.speech.with_streaming_response.create(
    model=os.environ["MODEL"],
    voice="alloy",
    input="The boat went over the weir.",
) as response:
    response.stream_to_file("speech.mp3")

with open("interview.mp3", "rb") as f:
    print(client.audio.transcriptions.create(model=os.environ["MODEL"], file=f).text)
```

```js
import { createWriteStream } from "node:fs";
import { Readable } from "node:stream";
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.HYPIT_API_KEY,
  baseURL: process.env.HYPIT_BASE_URL ?? "https://hypit.ai/v1",
});

const speech = await client.audio.speech.create({
  model: process.env.MODEL,
  voice: "alloy",
  input: "The boat went over the weir.",
});
Readable.fromWeb(speech.body).pipe(createWriteStream("speech.mp3"));
```

## Errors specific to these routes [#errors-specific-to-these-routes]

| Status | `code`                                                  | Route                  | Meaning                                          |
| ------ | ------------------------------------------------------- | ---------------------- | ------------------------------------------------ |
| `400`  | `missing_model`                                         | all                    | `model` is required                              |
| `400`  | `missing_input`                                         | speech, music          | no text, prompt, lyrics or voice description     |
| `413`  | `text_too_long`                                         | speech, music          | over 40 000 characters                           |
| `400`  | `invalid_n`                                             | speech, music          | `n` must be 1–8                                  |
| `400`  | `invalid_duration`                                      | speech, music          | duration must not be negative                    |
| `400`  | `invalid_response_format`                               | speech, music          | not one of the seven audio formats               |
| `400`  | `invalid_output`                                        | speech, music          | not `binary`, `b64_json` or `url`                |
| `413`  | `voice_sample_too_large`                                | speech                 | the MiMo VoiceClone base64 sample exceeds 10 MiB |
| `400`  | `unsupported_content_type`                              | transcriptions         | neither JSON nor multipart                       |
| `400`  | `missing_file` / `ambiguous_input`                      | transcriptions         | supply exactly one of `file` / `url`             |
| `400`  | `invalid_response_format`                               | transcriptions         | not `json`, `text` or `verbose_json`             |
| `400`  | `invalid_timestamp_granularity`                         | transcriptions         | not `segment` or `word`                          |
| `400`  | `invalid_temperature`                                   | transcriptions         | not a number                                     |
| `400`  | `unsupported_audio_format`                              | transcriptions         | the bytes are not a recognised container         |
| `400`  | `empty_file` / `unreadable_file`                        | transcriptions         | the part carried nothing usable                  |
| `413`  | `file_too_large`                                        | transcriptions         | over 100 MiB                                     |
| `413`  | `body_too_large`                                        | all                    | over the per-route body cap                      |
| `502`  | `empty_audio_response` / `empty_transcription_response` | speech, transcriptions | the upstream returned nothing usable             |