Skip to main content
Use this quickstart to analyze a recording with no picture, such as a phone call, a podcast, or a voice note, and get your first result in a few minutes. In this guide, you will:
  1. Submit a local audio file to POST /v2/upload/analyze with model=inter-2-audio, which creates an analysis job
  2. Poll the job until it finishes
  3. Read each window’s engagement_status and signals[] (including per-signal rationale), plus the optional whole-file signals list and conversation_quality output
You’ll need an API key. Follow the API key guide for details. You’ll also need an audio recording of at least 3 seconds in which people are speaking.

Inter-2 Audio

Inter-2 Audio (inter-2-audio) is the Inter-2 model that listens. It reads the sound of each analysis window alone, and reports the same kind of result as Inter-2: an engagement reading and the social signals it heard, each with a short rationale that cites what was audible. To analyze a video’s picture and sound together, use Inter-2 instead: see Video Analysis.

1) Submit an audio file

Send a local file to POST /v2/upload/analyze with model=inter-2-audio.

Supported files

The API identifies the format from the file’s contents, not from its name or MIME type. The samples below set a MIME type such as audio/wav by convention; use the one that matches your file (for example audio/mpeg, audio/flac, audio/mp4, or audio/ogg), or none. A file is refused before any analysis when:
  • It has no audio track: ih5008.
  • Its format isn’t one the model reads: ih4002. The error message lists the formats it does.
  • It can’t be read, for example a corrupted file: ih5002.
  • It’s empty or has no measurable duration: ih5001.
  • It’s under 3 seconds: ih4007. It’s over the duration limit: ih4004.
The job returns core analysis by default. You can optionally request whole-file sections by passing include[] flags (shown below):
  • signals adds one signal list over the whole file, with one entry per continuous span of each signal type.
  • conversation_quality_overall and conversation_quality_timeline add conversation-quality scores.
The API answers 202 with a job envelope. Its status is queued, and its status_url is the path to poll in the next step.
If you want only core outputs, remove the include[] lines. For another format, change the file name and the MIME type to match; the model stays inter-2-audio. The credential needs the interhumanai.upload scope, the same scope as any upload job. If the deployment serves no Inter-2 Audio backend at the moment, the submit answers 503 with ih1003; retry later.

2) Poll the job

Read GET /v2/upload/jobs/{job_id} until status is completed or failed. The status moves through queued and running first. status_url is a path relative to https://api.interhuman.ai.
To skip polling for short files, add wait_seconds (up to 120) to the submit. The API then holds the request open and answers 200 with the job if it reaches completed or failed in time, so check status either way. If it doesn’t finish in time, you get the usual 202 with the pending job, so keep the polling loop. A failed job carries an error in the standard error shape instead of a result. A job stays readable for 1 hour after submission; after that, reading it returns 404 (ih4021). Reference: Inter-2 and Inter-2 Audio Upload Analyze

3) Read the result

A completed job has the same shape for both Inter-2 models. Its result divides the recording into fixed-length windows and reports each one:
  • windows[]: one entry per window, in file order, with its span (start_seconds, end_seconds).
  • windows[].engagement_status: the attention level the speech conveys over that window: engaged, neutral, or disengaged.
  • windows[].signals[]: social signals heard in that window, each with type, probability, and rationale. A signal’s start and end are its window’s span.
  • signals (optional, with include[]=signals): every signal in the file as one list, in the order they start. A signal of the same type in adjacent windows, or reported more than once in a window, becomes one entry spanning them, with the highest probability and its rationale (on a tie, the earliest one’s). A window without that signal ends the entry. It is an empty list when nothing was detected.
  • conversation_quality (optional): computed once over the whole file. It uses the conversation_quality_values shape in both overall and each timeline entry’s values object.
Time fields are expressed in seconds from the start of the uploaded recording. The job envelope’s model is inter-2-audio.

The signals Inter-2 Audio reports

Inter-2 Audio concludes everything from what it hears, so it reports a subset of the signals Inter-2 reports. The signals whose evidence is mostly visual come from Inter-2 only: Every signal Inter-2 Audio reports is one Inter-2 can report too, so code written for Inter-2 results reads Inter-2 Audio results unchanged. The signals carry no modality field: the job’s model already says what was read. conversation_quality_values shape (used by conversation_quality.overall and conversation_quality.timeline[].values):
Here’s an example of a completed job:
probability and rationale are optional on every signal, so treat both as possibly absent.

How to interpret it quickly

  • windows[].signals[] gives the signals heard in each window; rationale explains each one from audible cues such as pauses, pitch, pace, and word choice.
  • signals gives the same signals over the whole file, one entry per continuous span, which is easier to show as a timeline.
  • windows[].engagement_status shows the attention level window by window.
  • conversation_quality.overall is a single interaction summary.
  • conversation_quality.timeline[] shows how quality changes over time.

Next steps