Skip to main content
The stream analysis quickstart shows how to call the stream endpoint. This guide covers what the quickstart deliberately leaves out: how to put that call inside a real application — one where the camera runs in a user’s browser and your API key stays secret. If you follow only three rules from this page, make them these:
  1. Never ship your API key to the browser. Mint a short-lived, capped client token server-side and hand that to the client instead.
  2. Let the browser stream directly to Interhuman with the client token. The token is designed for exactly this, and the API enforces its caps.
  3. Send media segments exactly as the recorder produces them. Never re-slice, reassemble, or reinterpret the media container in between.
This guide uses WS /v2/stream/analyze, analyzed by Inter-2, which requires sound as well as a picture: record with both the camera and the microphone. For an integration still on WS /v1/stream/analyze, see Migrate from V1 to V2.
This guide is written for humans and for AI coding agents. If you are prompting an agent to build an Interhuman integration, include this page (or paste the three rules above) — the naive implementation that hardcodes the API key in client code works in a demo and leaks your credential in production.

Why you can’t ship the API key

The stream endpoint authenticates with your API key (ih_live_...) as a Bearer credential on the WebSocket handshake. If that key appears anywhere in client code, it ships to the user — it’s in your JS bundle, visible in the network tab, extractable from anyone’s devtools. That is a leaked production credential.
Framework-specific version of the same leak: any environment variable prefixed NEXT_PUBLIC_ (Next.js), VITE_ (Vite), or REACT_APP_ (CRA) is bundled into the browser build. Never put your Interhuman API key in one of these.
The key must stay server-side. What the browser gets instead is a client token: a short-lived credential your backend mints from POST /v1/client_tokens, scoped and capped so that leaking it costs you a few minutes of bounded usage, not your account.

The architecture

Two parts, one job each:
  • A token route on your backend calls POST /v1/client_tokens with your API key and returns the minted token to the browser. This is stateless, single-request work — a serverless function is fine here.
  • The browser connects directly to wss://api.interhuman.ai/v2/stream/analyze with the client token, streams recorded segments up as binary frames, and receives analysis events back. The caps embedded in the token (duration, bytes, concurrency, video budget, allowed origins) are enforced by the API itself.
No media ever touches your servers, and there is no socket for you to host.

Mint a client token

Use the TypeScript SDK (@interhumanai/sdk) in your token route, or call the endpoint directly:
Every cap is optional, but set them deliberately — they are your blast radius if a token leaks: You can also revoke a token early (POST /v1/client_tokens/revoke, or auth.revokeClientToken(...) in the SDK) — new requests are rejected immediately and any live session is torn down on its next chunk. Useful when a token outlives the thing it was minted for, like a user logging out mid-session.

Connect from the browser

The SDK’s StreamClient handles the connection, authentication, typed events, and graceful shutdown:
Without the SDK: browsers cannot set an Authorization header on a WebSocket, so pass the token via the subprotocol pair the endpoint accepts:

Send media the way the recorder produces it

This is where integrations lose the most time. The producer of the media — MediaRecorder in the browser — already emits a well-formed media stream (WebM in most browsers) in pieces when you record with a timeslice. The first blob carries the stream’s header, and each later blob continues the same stream. Send every blob, in order, on the same session:
Each 3-second blob is small (roughly 400 KB at 1 Mbps) and picks up where the previous one ended. Send it the moment it is produced, and nothing downstream ever has to reverse-engineer the container structure. What does not work, in increasing order of subtlety:
  • One recording, one giant frame. The endpoint enforces a 32 MB maximum WebSocket message size; a 3-minute recording at typical bitrates exceeds it, and you get an ih6002 (“message too large”) error or the upstream resets the socket mid-send (close code 1006, no close frame). Worse, large single messages fail unreliably even under the cap. Small and many beats big and one.
  • Re-slicing the media in transit. Parsing the WebM/EBML structure yourself and cutting your own chunk boundaries is fragile: dropped trailing metadata or a mis-cut boundary produces a truncated container and an ih5004 (“malformed segment”) error. Container parsers that pass tests against synthetic ffmpeg files still mis-cut real live-muxed browser output.
  • Concatenating the blobs and sending the result as one frame. The joined file has no duration or seek index (MediaRecorder writes neither while recording), and as one frame it runs into the 32 MB cap above. Stream each blob live, in order, on one session instead.
  • Sending a later blob on its own. Only the first blob carries the header, so a blob sent without the ones before it, such as on a new session, can’t be decoded. Start a new session with a new recording.

End the session cleanly

Analysis events stream in throughout the recording, so most results have already arrived by the time the user stops. Accumulate events client-side as they come — there is no end-of-session summary to wait for. When the user stops recording, request a graceful shutdown instead of closing the socket. Stop the recorder, and let the dataavailable handler above call stream.requestClose() (which sends session.close) after it sends the final blob. The server refuses video that arrives after session.close, so closing first would drop the last segment.
The server acknowledges with session.closing (its data.max_drain_seconds is the longest you should wait), finishes analyzing the video it already accepted — emitting the remaining envelopes, plus signal.ended for still-active signals — then sends session.ended and closes with code 1000. Listen for session.ended (or close) rather than closing the socket yourself; closing early discards trailing analysis.

Debugging: symptom → cause

Next steps