Overview
The Audimee API lets you:- Browse Audimee’s built-in voice models and instruments, and your own custom-trained voices
- Train new custom voice models from your own audio
- Submit audio files for AI voice conversion against any kind of model
- Turn a single vocal into a multi-voice harmony
Base URL
Authentication
All endpoints require a static API key issued by Audimee, passed as a bearer token:401 with one of:
Missing or invalid Authorization header— header absent or not prefixed withBearerInvalid authorization token— token not recognisedUser not found— token valid but the underlying user no longer exists
Error responses
Every error response — across every endpoint — uses the same shape:400 validation, 401 auth, 403 forbidden, 404 not found, 500 server error). The error string is the specific reason.
Voice model types
Responses from/voice-models are a discriminated union on type:
audimee— Built-in voices. Always available, no training. IDs are short numeric strings like"101". Includes display fields (description,previewUrl,gender,genres,pitchLow,pitchHigh).instrument— Built-in instruments (guitars, saxophones, drums and more). Always available, no training. IDs are short numeric strings like"1013". CarrydescriptionandpreviewUrlonly — no gender, genres or pitch range. Work best with wordless vocal input such as hummed melodies or ad-libs that mimic the instrument.custom— Voices you trained viaPOST /voice-models. IDs are UUIDs. Become usable for conversions onceconfirmedDone: true. Carry training state (percentageDone,confirmedDone,failed,failedMessage).
?type=audimee, ?type=instrument or ?type=custom on GET /voice-models to narrow the response.
Custom voice model lifecycle
- Start training.
POST /voice-modelswithnameandtrainingAudioUrls. The response returns immediately with{ id, type: "custom" }. Training is queued on the AI backend. - Poll progress.
GET /voice-models/{id}returns the currentpercentageDone(0–100). Recommended cadence: every 10–30s — training typically takes minutes. - Detect terminal state. A model is terminal when one of:
confirmedDone: true, failed: false— ready for use in conversionsfailed: true— training failed;failedMessagecarries the reason
- Use it. Pass the custom model’s UUID as
voiceModelIdwhen callingPOST /conversions. - Delete it (optional).
DELETE /voice-models/{id}permanently deletes the model and its training artifacts — this is irreversible. The model must have started training on the AI backend (percentageDone > 0) before delete is allowed — if it hasn’t, the call returns400and you should retry later. If training is still in progress, the delete cancels it first.
Conversion lifecycle
- Start the conversion.
POST /conversionswithvoiceModelIdandinputFileUrl. Optional knobs:conversionStrength,pitchShift,sampleRateHz. Response:{ id }. - Poll progress.
GET /conversions/{id}returnspercentageDoneandconfirmedDone. Recommended cadence: every 5–15s — conversions are typically faster than training. - Detect terminal state.
confirmedDone: true(success) orfailed: true(failure). - Download the output. When
confirmedDoneistrue, theoutputFilePlayUrl(MP3, for streaming) andoutputFileDownloadUrl(WAV when available, MP3 otherwise) are usable signed URLs. The URLs are present in the response from the moment the conversion is created, but they don’t resolve to audio until the job is done — gate access onconfirmedDone, not on URL contents.
Harmony lifecycle
- Pick a preset.
GET /harmony-presetslists Audimee’s harmony presets — the same ones offered in the harmony maker. Each has anid, a displaynameand 1–5tracks. - Start the harmony.
POST /harmonieswithinputFileUrl, the musicalkeyof the input (for example"C Maj") and the preset’sidaspresetId. Response:{ id }. The call answers once the input’s pitch has been analysed (a few seconds, up to about a minute when that service is starting up), so use a client timeout of at least 120 seconds and don’t retry after a timeout. - Poll progress.
GET /harmonies/{id}returns every track with its ownpercentageDoneandconfirmedDone. Recommended cadence: every 5–15s. - Detect terminal state. Per track:
confirmedDone: true(success) orfailed: true(failure); one failed track does not stop the others. Tracks not finished after about 16 minutes are removed fromtracks, so comparetracks.lengthwith the preset; once every track is gone the harmony returns404. - Download and mix. Each track has its own
outputFilePlayUrlandoutputFileDownloadUrl— tracks are delivered separately, not mixed. Apply each track’spanandtimeShiftin your mix to reproduce the stereo spread of Audimee’s harmony maker.
Plan limits
API client accounts may have plan-driven limits:- Maximum concurrent custom models — when reached,
POST /voice-modelsreturns400withCustom model limit reached (max N). - Maximum combined training audio duration per model — default 30 minutes. When exceeded,
POST /voice-modelsreturns400with the actual vs. allowed seconds. - Harmony input length — at most 75 seconds. When exceeded,
POST /harmoniesreturns400with the actual vs. allowed seconds. - Conversion time — a harmony uses conversion time equal to the input length × the number of tracks in the preset. When the account doesn’t have enough left,
POST /harmoniesreturns400with the seconds needed.
trainingAudioUrls and inputFileUrl must be publicly reachable URLs — the server downloads them server-side. The host doesn’t need to be Audimee’s, but it does need to be available for the duration of the API call.