Send a recording
Send a recording and whatever you know about it. Caller creates the call, queues it for transcription and responds with its ID. One request per call.
Request
Section titled “Request”The body is multipart/form-data: a form with the file and, if you want, text fields. The only header you have to set is the authentication header. The library you use sets the Content-Type of the form.
curl https://api.caller.ee/v1/calls \ -H "Authorization: Bearer $CALLER_API_KEY" \ -F "file=@recordings/pbx-000123.mp3" \ -F "external_id=pbx-000123" \ -F "agent_code=AG1143" \ -F "campaign=Retention" \ -F "started_at=2026-10-09T10:42:11+02:00" \ -F "customer_phone=+34600000000" \ -F "direction=inbound"import osimport requests
with open("recordings/pbx-000123.mp3", "rb") as audio: response = requests.post( "https://api.caller.ee/v1/calls", headers={"Authorization": f"Bearer {os.environ['CALLER_API_KEY']}"}, files={"file": ("pbx-000123.mp3", audio, "audio/mpeg")}, data={ "external_id": "pbx-000123", "agent_code": "AG1143", "campaign": "Retention", "started_at": "2026-10-09T10:42:11+02:00", "customer_phone": "+34600000000", "direction": "inbound", }, timeout=120, )
body = response.json()if not response.ok: raise RuntimeError(f"{response.status_code} {body['error']['code']}")print(body["id"], body["status"])// Node.js 18 or later: fetch, FormData and Blob are built in.import { readFile } from 'node:fs/promises';
const audio = new Blob([await readFile('recordings/pbx-000123.mp3')], { type: 'audio/mpeg' });
const form = new FormData();form.append('file', audio, 'pbx-000123.mp3');form.append('external_id', 'pbx-000123');form.append('agent_code', 'AG1143');form.append('campaign', 'Retention');form.append('started_at', '2026-10-09T10:42:11+02:00');form.append('customer_phone', '+34600000000');form.append('direction', 'inbound');
const response = await fetch('https://api.caller.ee/v1/calls', { method: 'POST', headers: { Authorization: `Bearer ${process.env.CALLER_API_KEY}` }, body: form,});
const body = await response.json();if (!response.ok) throw new Error(`${response.status} ${body.error.code}`);console.log(body.id, body.status);<?php$request = curl_init('https://api.caller.ee/v1/calls');curl_setopt_array($request, [ CURLOPT_POST => true, CURLOPT_RETURNTRANSFER => true, CURLOPT_TIMEOUT => 120, CURLOPT_HTTPHEADER => ['Authorization: Bearer ' . getenv('CALLER_API_KEY')], CURLOPT_POSTFIELDS => [ 'file' => new CURLFile('recordings/pbx-000123.mp3', 'audio/mpeg', 'pbx-000123.mp3'), 'external_id' => 'pbx-000123', 'agent_code' => 'AG1143', 'campaign' => 'Retention', 'started_at' => '2026-10-09T10:42:11+02:00', 'customer_phone' => '+34600000000', 'direction' => 'inbound', ],]);
$body = json_decode(curl_exec($request), true);$status = curl_getinfo($request, CURLINFO_RESPONSE_CODE);curl_close($request);
if ($status >= 400) { throw new RuntimeException($status . ' ' . $body['error']['code']);}echo $body['id'], ' ', $body['status'], PHP_EOL;Fields
Section titled “Fields”Only file is required. Text fields are truncated to 300 characters.
| Field | Type | Description |
|---|---|---|
file |
file | Required. The recording. See formats. |
external_id |
text | The ID of the call in your system, up to 200 characters. Sending the same ID twice does not duplicate the call. See retries. |
agent_code |
text | The agent’s code in your telephony system. It is compared with the Code in the telephony system of each agent in Team, ignoring case, hyphens, underscores and spaces. |
agent_name |
text | The agent’s name, as it is in Team. It is used if there is no code or if no agent has that code. |
campaign |
text | The name of a campaign that already exists in the account. Case and accents are ignored. |
started_at |
date | When the call started: ISO 8601 with its UTC offset (2026-10-09T10:42:11+02:00), or Unix time in seconds or in milliseconds. |
customer_phone |
text | The customer’s phone number. |
customer_name |
text | The customer’s name. |
direction |
text | inbound if the call was incoming, outbound if it was outgoing. |
language |
text | The language code of the call, so that it is not left to automatic detection. See language. |
What you do not send is read from the file name
Section titled “What you do not send is read from the file name”If the agent, the date, the phone number or the direction is missing, Caller tries to read it from the file name, as it does when you upload by hand. A field you send always wins over the file name. The rules are in File names.
The agent is required
Section titled “The agent is required”Everything else is optional, but every recording has to say whose call it is, and that person has to be in Team already. Any of these three ways will do: agent_code, agent_name, or the agent’s code in the file name.
- If the recording does not say whose it is, the response is
422withagent_required. - If it does but nobody in Team has that code or that name, the response is
422withunknown_agent.
In both cases the recording is not kept and uses no minutes. Agents are added in the app, under Team, with the same code your telephony system uses; each one takes one of your plan’s seats.
With a telephony system that already names its recordings well, the minimal request is the file and its external_id.
When a field is ignored
Section titled “When a field is ignored”Except for the agent, which is required, the API does not reject a request because of a value it does not recognize: it creates the call without that value. This is good to know, because you will not see an error.
| Situation | What happens |
|---|---|
campaign does not exist or is archived |
The call comes in without a campaign and is scored with the default template. Campaigns are not created automatically. |
started_at cannot be read as a date |
It is ignored. If the file name has a date, that date is used. If not, the call is dated at the moment of the upload. |
started_at is earlier than the year 2000 or more than one day in the future |
It is ignored, and in this case the file name is not looked at either: the call is dated at the moment of the upload. |
direction has any other value |
It is ignored. |
Response
Section titled “Response”201 Created when the call has been created:
{ "id": "k57e2xq9m4hc8w1t6b0z5y2e97c4n1ad", "status": "queued" }200 OK when a call with that external_id already existed. Nothing new is created and the file you sent is discarded:
{ "id": "k57e2xq9m4hc8w1t6b0z5y2e97c4n1ad", "status": "exists" }| Field | Description |
|---|---|
id |
The ID of the call in Caller. It is the same one that appears in the webhook and in the address of the call in the app: https://app.caller.ee/calls/{id}. |
status |
queued if the call is new and waiting in the queue; exists if it already existed. |
Any other status code is an error, with this body:
{ "error": { "code": "not_audio", "message": "The file is not an audio recording." } }The full list is in Errors.
Retry without duplicates
Section titled “Retry without duplicates”Networks fail. If a request is cut off and you do not know whether it arrived, send it again with the same external_id:
- If the first request did not arrive, the call is created and you get
201. - If it did arrive, you get
200with"status": "exists"and theidof the call that was already there.
Either way, nothing is charged twice. Use the unique ID your telephony system gives each call as the external_id.
Without external_id, every request creates a new call, even if the file is identical.
Formats and size
Section titled “Formats and size”- Formats: MP3, WAV, OGG and M4A are the usual ones. AAC, FLAC, Opus, WebM, WMA and AMR are also accepted. If your telephony system records in one of the latter formats, test with a few calls before you automate sending.
- Maximum size: the whole request, with the file and the fields, has to be under 20 MB. This is a limit of the API, lower than the per-file maximum of some plans: a recording between 20 MB and your plan’s maximum can be uploaded from the app, but not through the API.
- If you go over: a request of 20 MB or more is rejected with
400 bad_request, the same code as a malformed body. Check the size before sending and compress what does not fit: for speech, mono MP3 at 64 kbps takes about half a megabyte per minute. - File type: if your system sends the file as
application/octet-stream, Caller goes by the extension of the file name. A file with no audio extension and noaudio/…type is rejected withnot_audio.
Language
Section titled “Language”Without language, the language of the campaign is used and, if the campaign has none, the language is detected for each call. Set it when you know it for certain: detection can confuse similar languages and needs about a minute of conversation to get it right.
Supported codes:
ca es en pt gl eu fr de it nl ar zh ru ja ko pl ro tr uk sv no da fi cs el hu he hi id th vi bg hr sk sr fa ur ms ta
What each request uses
Section titled “What each request uses”Each accepted call takes up its size in storage and reserves minutes of audio. Its analysis is included in those minutes.
The reservation is based on the size of the file, at one minute per megabyte, and is adjusted to the real duration when the transcription finishes. With MP3, the reservation is close to the duration. With uncompressed WAV it can be several times larger, and a recording can be rejected with no_minutes even though its real duration would have fit. If you are short on minutes, send compressed audio.
What happens next
Section titled “What happens next”The call follows the same path as a manual upload: it is transcribed, analyzed and shown in the app. Once it is analyzed, Caller sends the result to your webhook, if you have one.
There is no request to check the status of a call. If a recording cannot be transcribed, you see it in the app, in the Failed view. No notice is sent.