Skip to main content
Stream analysis from a live camera: connect to Interhuman over WebSocket, capture media, and send binary chunks as they are recorded. In this codealong, you will:
  1. Connect to wss://api.interhuman.ai/v2/stream/analyze
  2. Get camera and microphone
  3. Send each recorded segment to Interhuman and read typed server events
Wire the three steps together in your app. Use JavaScript in the browser (getUserMedia + MediaRecorder) or Python on the desktop (ffmpeg for camera and microphone capture). Sessions are analyzed by Inter-2, which requires sound as well as a picture. Every segment you send must carry an audio track as well as video.
You’ll need an API key. Follow the API key guide for details. In the browser, use a short-lived client token minted from that key on your server instead of the key itself.
Already using WS /v1/stream/analyze? See Migrate from V1 to V2.

1) Connect to the WebSocket

Open a TLS WebSocket to the stream endpoint. On connect, send session config as a text frame (UTF-8 JSON), then listen for text replies and parse JSON. Branch on type. The analysis events are signal.detected, signal.updated, signal.ended, engagement.updated, conversation_quality.updated, and error. The connection also sends session and coverage events (session.ready, session.updated, session.closing, session.ended, coverage.degraded, coverage.dropped), listed in Inter-2 Streaming Analyze. The session uses the inter-2 model by default, so the session config doesn’t need to name one. If no backend is serving the model, the server refuses the session after the handshake with an ih1003 error envelope and close code 1013; retry later.
Reference: Inter-2 Streaming Analyze

2) Get camera and microphone

Device names and indexes vary by machine. To list them, run ffmpeg -f avfoundation -list_devices true -i "" on macOS or v4l2-ctl --list-devices on Linux. On macOS, the terminal also needs camera and microphone permission. On Linux, the pulse input needs an ffmpeg build with PulseAudio support.

3) Send segments to Interhuman

Send the recording as one continuous media stream, split into binary WebSocket frames. The first frame carries the container’s header, and each later frame continues the same stream: WebM clusters from MediaRecorder, or fragmented-MP4 fragments from ffmpeg. Don’t send each segment as a separate file. Start recording after the WebSocket is open and session config is sent.
When the user stops, stop recording and ask the server to end the session with a session.close text frame. Don’t close the socket yourself. The server finishes analyzing the video it already accepted, sends the remaining envelopes (including signal.ended for signals still active), sends session.ended, and then closes the connection. If you close the socket early, those last envelopes are lost. See End the session cleanly.

4) Read server envelopes

Every server message shares the same outer shape: type, timestamp, correlation_id, and data. Narrow on type before reading fields inside data. This section covers the analysis events; the session and coverage events are in Inter-2 Streaming Analyze.

signal.detected

Each signal.detected envelope reports one signal becoming active. It always carries data.signal_type and data.start. It may also carry data.probability and data.rationale, so treat those two as optional. It carries no end time.

signal.updated and signal.ended

While a signal stays active, signal.updated reports a change to its probability or rationale, with the same data fields as signal.detected. On signal.updated, data.start is when the new state applies, not when the signal first became active; keep the start from signal.detected to measure the signal’s full span. When the signal is no longer active, signal.ended reports its data.signal_type and data.end:

engagement.updated

engagement.updated is sent when engagement is first established and again each time the level changes. data.state is engaged, neutral, or disengaged, and it holds until the next engagement.updated.

conversation_quality.updated (when opted in)

When your session config include lists conversation_quality_overall and/or conversation_quality_timeline, you may receive conversation_quality.updated with data.overall and/or data.timeline for the window that was just processed. See Conversation quality.

error

Errors use the same envelope with type: "error" and structured fields under data (for example code, message, link, and segment when applicable). See Error handling.

How to interpret it quickly

  • signal.detected, signal.updated, signal.ended: the lifecycle of one social signal, from the moment it becomes active until it ends. A rationale, when present, explains the detection.
  • engagement.updated: the attention level from start onward, until the next engagement.updated.
  • Times: start and end are in seconds, cumulative across the whole session rather than relative to each segment.
  • conversation_quality.updated: optional overall and per-window quality metrics when requested.

Next steps