Skip to content
All projects

Timbre

AI video soundtracking platform

Analyzes a video’s scenes and dialogue, then streams a score that follows its mood in real time.

Status
Shipped
Period
Sep 2025 - Oct 2025
Built with
  • FastAPI
  • WebSockets
  • Groq
  • Google Lyria
  • React
  • Next.js

What it does

  • Built a video-to-music system that detects scenes, samples each scene’s middle frame, and pairs those images with a transcript for multimodal analysis.
  • Used Groq and a Llama model to turn the video analysis into a music prompt and mood direction for every scene.
  • Connected the client, backend, and real-time music API with WebSockets, streaming a new audio chunk every two seconds and handling buffering for continuous playback.
  • Built the live experience end to end, coordinating scene analysis, transcription, prompt generation, streaming, buffering, and playback at sub-two-second latency.

Why I built it

A score can completely change how a video feels. When I saw that real-time music APIs made immediate generation possible, I wanted to make it possible for any video to receive a background score without treating music as a slow export step.

How it works

Timbre detects scenes and analyzes the transcript in parallel. It selects a representative frame from every scene, then sends those images and the transcript to Groq and a Llama model to determine the prompt and mood for the music. The client, backend, and real-time music API stay connected over WebSockets while the score streams back in two-second chunks.

The hard part

The challenge was connecting the moving parts into one coherent experience. Scene detection, transcription, multimodal prompting, streaming, buffering, and playback all had to work together closely enough that the user hears a continuous score instead of a collection of separate systems.

Proof in the scene

The moment it clicked was using an Avengers: Infinity War fight scene and hearing Timbre generate a score that matched the action. Timbre is live today, so the project is more than a technical pipeline: it is an experience people can try.

How it works

  1. Video

    uploaded by the viewer

  2. Scene + transcript

    representative frames and dialogue

  3. Musical plan

    Groq and Llama set the prompt and mood

    new audio chunk every 2 s
  4. Live score

    streamed over WebSockets, buffered for playback

Timbre is live at timbreapp.tech, with sub-two-second end-to-end latency.

Next projectGobble