> ## Documentation Index
> Fetch the complete documentation index at: https://docs.interhuman.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio Analysis

> Submit an audio recording to Inter-2 Audio, poll the job, and read the signals it heard.

Use this quickstart to analyze a recording with no picture, such as a phone call, a podcast, or a voice note, and get your first result in a few minutes.

In this guide, you will:

1. Submit a local audio file to `POST /v2/upload/analyze` with `model=inter-2-audio`, which creates an analysis job
2. Poll the job until it finishes
3. Read each window's `engagement_status` and `signals[]` (including per-signal `rationale`), plus the optional whole-file `signals` list and `conversation_quality` output

<Note>
  You’ll need an API key. Follow the [API key guide](/how-to/get-api-key) for details.
  You’ll also need an audio recording of at least 3 seconds in which people are speaking.
</Note>

## Inter-2 Audio

**Inter-2 Audio** (`inter-2-audio`) is the Inter-2 model that listens. It reads the sound of each analysis window alone, and reports the same kind of result as Inter-2: an engagement reading and the social signals it heard, each with a short rationale that cites what was audible.

To analyze a video's picture and sound together, use Inter-2 instead: see [Video Analysis](/getting-started/video-upload-quickstart).

## 1) Submit an audio file

Send a local file to `POST /v2/upload/analyze` with `model=inter-2-audio`.

### Supported files

| Requirement | Value |
| - | - |
| Formats | wav, flac, mp3, m4a, or ogg, or a webm or mp4 that carries an audio track |
| Tracks | At least one audio track. A video track, if present, is ignored. |
| Duration | At least 3 seconds, and no longer than 30 minutes |
| Size | At most 32 MB |

The API identifies the format from the file's contents, not from its name or MIME type. The samples below set a MIME type such as `audio/wav` by convention; use the one that matches your file (for example `audio/mpeg`, `audio/flac`, `audio/mp4`, or `audio/ogg`), or none.

A file is refused before any analysis when:

* It has no audio track: `ih5008`.
* Its format isn't one the model reads: `ih4002`. The error message lists the formats it does.
* It can't be read, for example a corrupted file: `ih5002`.
* It's empty or has no measurable duration: `ih5001`.
* It's under 3 seconds: `ih4007`. It's over the duration limit: `ih4004`.

The job returns core analysis by default. You can optionally request whole-file sections by passing `include[]` flags (shown below):

* `signals` adds one signal list over the whole file, with one entry per continuous span of each signal type.
* `conversation_quality_overall` and `conversation_quality_timeline` add conversation-quality scores.

The API answers `202` with a **job envelope**. Its `status` is `queued`, and its `status_url` is the path to poll in the next step.

<CodeGroup>
  ```bash cURL theme={null}
  export API_KEY="YOUR_API_KEY"
  export AUDIO_PATH="path_to_your_recording.wav"

  curl -X POST https://api.interhuman.ai/v2/upload/analyze \
    -H "Authorization: Bearer ${API_KEY}" \
    -F "file=@${AUDIO_PATH};type=audio/wav" \
    -F "model=inter-2-audio" \
    -F "include[]=signals" \
    -F "include[]=conversation_quality_overall" \
    -F "include[]=conversation_quality_timeline"
  ```

  ```javascript JavaScript icon="square-js" theme={null}
  // Run as an ES module (for example, save as quickstart.mjs) for top-level await.
  import fs from "fs";
  import FormData from "form-data";
  import fetch from "node-fetch";

  const formData = new FormData();
  formData.append("file", fs.createReadStream(process.env.AUDIO_PATH));
  formData.append("model", "inter-2-audio");
  formData.append("include[]", "signals");
  formData.append("include[]", "conversation_quality_overall");
  formData.append("include[]", "conversation_quality_timeline");

  const response = await fetch("https://api.interhuman.ai/v2/upload/analyze", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.API_KEY}`
    },
    body: formData
  });
  if (!response.ok) {
    throw new Error(`${response.status}: ${await response.text()}`);
  }
  let job = await response.json();
  console.log("Job:", job.job_id, job.status);
  ```

  ```python Python icon="python" theme={null}
  import os
  import requests

  api_key = os.environ["API_KEY"]
  audio_path = os.environ["AUDIO_PATH"]

  with open(audio_path, "rb") as f:
      files = {"file": (os.path.basename(audio_path), f, "audio/wav")}
      data = [
          ("model", "inter-2-audio"),
          ("include[]", "signals"),
          ("include[]", "conversation_quality_overall"),
          ("include[]", "conversation_quality_timeline"),
      ]
      response = requests.post(
          "https://api.interhuman.ai/v2/upload/analyze",
          headers={"Authorization": f"Bearer {api_key}"},
          files=files,
          data=data,
          timeout=300,
      )

  response.raise_for_status()
  job = response.json()
  print("Job:", job["job_id"], job["status"])
  ```
</CodeGroup>

If you want only core outputs, remove the `include[]` lines. For another format, change the file name and the MIME type to match; the `model` stays `inter-2-audio`.

The credential needs the `interhumanai.upload` scope, the same scope as any upload job. If the deployment serves no Inter-2 Audio backend at the moment, the submit answers `503` with `ih1003`; retry later.

## 2) Poll the job

Read `GET /v2/upload/jobs/{job_id}` until `status` is `completed` or `failed`. The status moves through `queued` and `running` first. `status_url` is a path relative to `https://api.interhuman.ai`.

<CodeGroup>
  ```bash cURL theme={null}
  # Paste the job_id from the submit response.
  export JOB_ID="YOUR_JOB_ID"

  curl https://api.interhuman.ai/v2/upload/jobs/${JOB_ID} \
    -H "Authorization: Bearer ${API_KEY}"
  ```

  ```javascript JavaScript icon="square-js" theme={null}
  while (job.status !== "completed" && job.status !== "failed") {
    await new Promise((resolve) => setTimeout(resolve, 2000));
    const res = await fetch(`https://api.interhuman.ai${job.status_url}`, {
      headers: { Authorization: `Bearer ${process.env.API_KEY}` }
    });
    if (res.status === 503) continue; // ih1002 / ih1003: retry after a short delay
    if (!res.ok) {
      throw new Error(`${res.status}: ${await res.text()}`);
    }
    job = await res.json();
  }
  console.log(JSON.stringify(job, null, 2));
  ```

  ```python Python icon="python" theme={null}
  import time

  while job["status"] not in ("completed", "failed"):
      time.sleep(2)
      response = requests.get(
          f"https://api.interhuman.ai{job['status_url']}",
          headers={"Authorization": f"Bearer {api_key}"},
          timeout=30,
      )
      if response.status_code == 503:
          continue  # ih1002 / ih1003: retry after a short delay
      response.raise_for_status()
      job = response.json()

  print(job)
  ```
</CodeGroup>

To skip polling for short files, add `wait_seconds` (up to 120) to the submit. The API then holds the request open and answers `200` with the job if it reaches `completed` or `failed` in time, so check `status` either way. If it doesn't finish in time, you get the usual `202` with the pending job, so keep the polling loop.

A `failed` job carries an `error` in the standard [error shape](/api-reference/error-handling) instead of a `result`. A job stays readable for 1 hour after submission; after that, reading it returns `404` (`ih4021`).

Reference: [Inter-2 and Inter-2 Audio Upload Analyze](/api-reference/upload-analyze-v2)

## 3) Read the result

A completed job has the same shape for both Inter-2 models. Its `result` divides the recording into fixed-length windows and reports each one:

* **windows\[]**: one entry per window, in file order, with its span (`start_seconds`, `end_seconds`).
* **windows\[].engagement\_status**: the attention level the speech conveys over that window: `engaged`, `neutral`, or `disengaged`.
* **windows\[].signals\[]**: [social signals](/explanations/social-signals) heard in that window, each with `type`, `probability`, and `rationale`. A signal's `start` and `end` are its window's span.
* **signals** (optional, with `include[]=signals`): every signal in the file as one list, in the order they start. A signal of the same type in adjacent windows, or reported more than once in a window, becomes one entry spanning them, with the highest `probability` and its `rationale` (on a tie, the earliest one's). A window without that signal ends the entry. It is an empty list when nothing was detected.
* **conversation\_quality** (optional): computed once over the whole file. It uses the `conversation_quality_values` shape in both `overall` and each timeline entry's `values` object.

Time fields are expressed in seconds from the start of the uploaded recording. The job envelope's `model` is `inter-2-audio`.

### The signals Inter-2 Audio reports

Inter-2 Audio concludes everything from what it hears, so it reports a **subset** of the signals Inter-2 reports. The signals whose evidence is mostly visual come from Inter-2 only:

| Reported by `inter-2-audio` | Reported by `inter-2` only |
| - | - |
| `agreement`, `confidence`, `frustration`, `hesitation`, `interest`, `uncertainty` | `confusion`, `disagreement`, `skepticism`, `stress` |

Every signal Inter-2 Audio reports is one Inter-2 can report too, so code written for Inter-2 results reads Inter-2 Audio results unchanged. The signals carry no `modality` field: the job's `model` already says what was read.

`conversation_quality_values` shape (used by `conversation_quality.overall` and `conversation_quality.timeline[].values`):

```json theme={null}
{
  "quality_index": 57.6,
  "clarity": 18.9,
  "authority": 18.9,
  "energy": 78.8,
  "rapport": 85.7,
  "learning": 85.7
}
```

Here’s an example of a completed job:

```json theme={null}
{
  "job_id": "8a2d4e6f0b1c4d3e9f8a7b6c5d4e3f21",
  "status": "completed",
  "model": "inter-2-audio",
  "created_at": "2026-10-07T10:00:00Z",
  "expires_at": "2026-10-07T11:00:00Z",
  "status_url": "/v2/upload/jobs/8a2d4e6f0b1c4d3e9f8a7b6c5d4e3f21",
  "result": {
    "duration_seconds": 10.0,
    "window_seconds": 5.0,
    "windows": [
      {
        "index": 0,
        "start_seconds": 0.0,
        "end_seconds": 5.0,
        "engagement_status": "engaged",
        "signals": [
          {
            "type": "hesitation",
            "start": 0.0,
            "end": 5.0,
            "probability": "medium",
            "rationale": "Long pauses and filler words (\"um, well...\") before answering, with a trailing, uncertain intonation."
          }
        ]
      },
      {
        "index": 1,
        "start_seconds": 5.0,
        "end_seconds": 10.0,
        "engagement_status": "engaged",
        "signals": [
          {
            "type": "interest",
            "start": 5.0,
            "end": 10.0,
            "probability": "high",
            "rationale": "Quick, animated replies with rising pitch and follow-up questions about the details."
          }
        ]
      }
    ],
    "signals": [
      {
        "type": "hesitation",
        "start": 0.0,
        "end": 5.0,
        "probability": "medium",
        "rationale": "Long pauses and filler words (\"um, well...\") before answering, with a trailing, uncertain intonation."
      },
      {
        "type": "interest",
        "start": 5.0,
        "end": 10.0,
        "probability": "high",
        "rationale": "Quick, animated replies with rising pitch and follow-up questions about the details."
      }
    ],
    "conversation_quality": {
      "overall": {
        "quality_index": 57.6,
        "clarity": 18.9,
        "authority": 18.9,
        "energy": 78.8,
        "rapport": 85.7,
        "learning": 85.7
      },
      "timeline": [
        {
          "start": 0.0,
          "end": 10.0,
          "values": {
            "quality_index": 57.6,
            "clarity": 18.9,
            "authority": 18.9,
            "energy": 78.8,
            "rapport": 85.7,
            "learning": 85.7
          }
        }
      ]
    }
  }
}
```

`probability` and `rationale` are optional on every signal, so treat both as possibly absent.

### How to interpret it quickly

* `windows[].signals[]` gives the signals heard in each window; `rationale` explains each one from audible cues such as pauses, pitch, pace, and word choice.
* `signals` gives the same signals over the whole file, one entry per continuous span, which is easier to show as a timeline.
* `windows[].engagement_status` shows the attention level window by window.
* `conversation_quality.overall` is a single interaction summary.
* `conversation_quality.timeline[]` shows how quality changes over time.

## Next steps

* [Stream analysis](/getting-started/stream-analyze-quickstart) — analyze a live microphone with Inter-2 Audio.
* [Social signals](/explanations/social-signals) — meaning of each signal type.
* [Conversation quality](/explanations/conversation-quality) — quality dimensions and interpretation.
* [Inter-2 and Inter-2 Audio Upload Analyze](/api-reference/upload-analyze-v2) — full request/response and error details.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.