- Submit a local audio file to
POST /v2/upload/analyzewithmodel=inter-2-audio, which creates an analysis job - Poll the job until it finishes
- Read each window’s
engagement_statusandsignals[](including per-signalrationale), plus the optional whole-filesignalslist andconversation_qualityoutput
You’ll need an API key. Follow the API key guide for details.
You’ll also need an audio recording of at least 3 seconds in which people are speaking.
Inter-2 Audio
Inter-2 Audio (inter-2-audio) is the Inter-2 model that listens. It reads the sound of each analysis window alone, and reports the same kind of result as Inter-2: an engagement reading and the social signals it heard, each with a short rationale that cites what was audible.
To analyze a video’s picture and sound together, use Inter-2 instead: see Video Analysis.
1) Submit an audio file
Send a local file toPOST /v2/upload/analyze with model=inter-2-audio.
Supported files
The API identifies the format from the file’s contents, not from its name or MIME type. The samples below set a MIME type such as
audio/wav by convention; use the one that matches your file (for example audio/mpeg, audio/flac, audio/mp4, or audio/ogg), or none.
A file is refused before any analysis when:
- It has no audio track:
ih5008. - Its format isn’t one the model reads:
ih4002. The error message lists the formats it does. - It can’t be read, for example a corrupted file:
ih5002. - It’s empty or has no measurable duration:
ih5001. - It’s under 3 seconds:
ih4007. It’s over the duration limit:ih4004.
include[] flags (shown below):
signalsadds one signal list over the whole file, with one entry per continuous span of each signal type.conversation_quality_overallandconversation_quality_timelineadd conversation-quality scores.
202 with a job envelope. Its status is queued, and its status_url is the path to poll in the next step.
include[] lines. For another format, change the file name and the MIME type to match; the model stays inter-2-audio.
The credential needs the interhumanai.upload scope, the same scope as any upload job. If the deployment serves no Inter-2 Audio backend at the moment, the submit answers 503 with ih1003; retry later.
2) Poll the job
ReadGET /v2/upload/jobs/{job_id} until status is completed or failed. The status moves through queued and running first. status_url is a path relative to https://api.interhuman.ai.
wait_seconds (up to 120) to the submit. The API then holds the request open and answers 200 with the job if it reaches completed or failed in time, so check status either way. If it doesn’t finish in time, you get the usual 202 with the pending job, so keep the polling loop.
A failed job carries an error in the standard error shape instead of a result. A job stays readable for 1 hour after submission; after that, reading it returns 404 (ih4021).
Reference: Inter-2 and Inter-2 Audio Upload Analyze
3) Read the result
A completed job has the same shape for both Inter-2 models. Itsresult divides the recording into fixed-length windows and reports each one:
- windows[]: one entry per window, in file order, with its span (
start_seconds,end_seconds). - windows[].engagement_status: the attention level the speech conveys over that window:
engaged,neutral, ordisengaged. - windows[].signals[]: social signals heard in that window, each with
type,probability, andrationale. A signal’sstartandendare its window’s span. - signals (optional, with
include[]=signals): every signal in the file as one list, in the order they start. A signal of the same type in adjacent windows, or reported more than once in a window, becomes one entry spanning them, with the highestprobabilityand itsrationale(on a tie, the earliest one’s). A window without that signal ends the entry. It is an empty list when nothing was detected. - conversation_quality (optional): computed once over the whole file. It uses the
conversation_quality_valuesshape in bothoveralland each timeline entry’svaluesobject.
model is inter-2-audio.
The signals Inter-2 Audio reports
Inter-2 Audio concludes everything from what it hears, so it reports a subset of the signals Inter-2 reports. The signals whose evidence is mostly visual come from Inter-2 only:
Every signal Inter-2 Audio reports is one Inter-2 can report too, so code written for Inter-2 results reads Inter-2 Audio results unchanged. The signals carry no
modality field: the job’s model already says what was read.
conversation_quality_values shape (used by conversation_quality.overall and conversation_quality.timeline[].values):
probability and rationale are optional on every signal, so treat both as possibly absent.
How to interpret it quickly
windows[].signals[]gives the signals heard in each window;rationaleexplains each one from audible cues such as pauses, pitch, pace, and word choice.signalsgives the same signals over the whole file, one entry per continuous span, which is easier to show as a timeline.windows[].engagement_statusshows the attention level window by window.conversation_quality.overallis a single interaction summary.conversation_quality.timeline[]shows how quality changes over time.
Next steps
- Stream analysis — analyze a live microphone with Inter-2 Audio.
- Social signals — meaning of each signal type.
- Conversation quality — quality dimensions and interpretation.
- Inter-2 and Inter-2 Audio Upload Analyze — full request/response and error details.