Skip to main content

Overview

The Audimee API lets you:
  • Browse Audimee’s built-in voice models and instruments, and your own custom-trained voices
  • Train new custom voice models from your own audio
  • Submit audio files for AI voice conversion against any kind of model
  • Turn a single vocal into a multi-voice harmony
Training, conversions and harmonies all run asynchronously — you start the job, then poll for its status.

Base URL

Authentication

All endpoints require a static API key issued by Audimee, passed as a bearer token:
Tokens are bound to a single API client account. Unauthorized requests return 401 with one of:
  • Missing or invalid Authorization header — header absent or not prefixed with Bearer
  • Invalid authorization token — token not recognised
  • User not found — token valid but the underlying user no longer exists

Error responses

Every error response — across every endpoint — uses the same shape:
The HTTP status code carries the category (400 validation, 401 auth, 403 forbidden, 404 not found, 500 server error). The error string is the specific reason.

Voice model types

Responses from /voice-models are a discriminated union on type:
  • audimee — Built-in voices. Always available, no training. IDs are short numeric strings like "101". Includes display fields (description, previewUrl, gender, genres, pitchLow, pitchHigh).
  • instrument — Built-in instruments (guitars, saxophones, drums and more). Always available, no training. IDs are short numeric strings like "1013". Carry description and previewUrl only — no gender, genres or pitch range. Work best with wordless vocal input such as hummed melodies or ad-libs that mimic the instrument.
  • custom — Voices you trained via POST /voice-models. IDs are UUIDs. Become usable for conversions once confirmedDone: true. Carry training state (percentageDone, confirmedDone, failed, failedMessage).
Use ?type=audimee, ?type=instrument or ?type=custom on GET /voice-models to narrow the response.

Custom voice model lifecycle

  1. Start training. POST /voice-models with name and trainingAudioUrls. The response returns immediately with { id, type: "custom" }. Training is queued on the AI backend.
  2. Poll progress. GET /voice-models/{id} returns the current percentageDone (0–100). Recommended cadence: every 10–30s — training typically takes minutes.
  3. Detect terminal state. A model is terminal when one of:
    • confirmedDone: true, failed: false — ready for use in conversions
    • failed: true — training failed; failedMessage carries the reason
  4. Use it. Pass the custom model’s UUID as voiceModelId when calling POST /conversions.
  5. Delete it (optional). DELETE /voice-models/{id} permanently deletes the model and its training artifacts — this is irreversible. The model must have started training on the AI backend (percentageDone > 0) before delete is allowed — if it hasn’t, the call returns 400 and you should retry later. If training is still in progress, the delete cancels it first.

Conversion lifecycle

  1. Start the conversion. POST /conversions with voiceModelId and inputFileUrl. Optional knobs: conversionStrength, pitchShift, sampleRateHz. Response: { id }.
  2. Poll progress. GET /conversions/{id} returns percentageDone and confirmedDone. Recommended cadence: every 5–15s — conversions are typically faster than training.
  3. Detect terminal state. confirmedDone: true (success) or failed: true (failure).
  4. Download the output. When confirmedDone is true, the outputFilePlayUrl (MP3, for streaming) and outputFileDownloadUrl (WAV when available, MP3 otherwise) are usable signed URLs. The URLs are present in the response from the moment the conversion is created, but they don’t resolve to audio until the job is done — gate access on confirmedDone, not on URL contents.

Harmony lifecycle

A harmony converts one vocal once per track of a harmony preset, each track singing its own line.
  1. Pick a preset. GET /harmony-presets lists Audimee’s harmony presets — the same ones offered in the harmony maker. Each has an id, a display name and 1–5 tracks.
  2. Start the harmony. POST /harmonies with inputFileUrl, the musical key of the input (for example "C Maj") and the preset’s id as presetId. Response: { id }. The call answers once the input’s pitch has been analysed (a few seconds, up to about a minute when that service is starting up), so use a client timeout of at least 120 seconds and don’t retry after a timeout.
  3. Poll progress. GET /harmonies/{id} returns every track with its own percentageDone and confirmedDone. Recommended cadence: every 5–15s.
  4. Detect terminal state. Per track: confirmedDone: true (success) or failed: true (failure); one failed track does not stop the others. Tracks not finished after about 16 minutes are removed from tracks, so compare tracks.length with the preset; once every track is gone the harmony returns 404.
  5. Download and mix. Each track has its own outputFilePlayUrl and outputFileDownloadUrl — tracks are delivered separately, not mixed. Apply each track’s pan and timeShift in your mix to reproduce the stereo spread of Audimee’s harmony maker.

Plan limits

API client accounts may have plan-driven limits:
  • Maximum concurrent custom models — when reached, POST /voice-models returns 400 with Custom model limit reached (max N).
  • Maximum combined training audio duration per model — default 30 minutes. When exceeded, POST /voice-models returns 400 with the actual vs. allowed seconds.
  • Harmony input length — at most 75 seconds. When exceeded, POST /harmonies returns 400 with the actual vs. allowed seconds.
  • Conversion time — a harmony uses conversion time equal to the input length × the number of tracks in the preset. When the account doesn’t have enough left, POST /harmonies returns 400 with the seconds needed.
Audio files passed via trainingAudioUrls and inputFileUrl must be publicly reachable URLs — the server downloads them server-side. The host doesn’t need to be Audimee’s, but it does need to be available for the duration of the API call.

Quick example — convert with a built-in voice

Quick example — train and use a custom voice

Quick example — create a harmony from a preset