Skip to content
All projects

EveryGPU

Distributed LLM inference engine

An experiment in joining spare GPUs into an efficient inference pipeline, with instrumentation that shows where every request waits.

Status
In progress
Period
Sep 2026 - Present
Built with
  • Python
  • CUDA
  • TCP
  • OpenTelemetry

What it does

  • Built a distributed inference prototype that passes activation vectors over TCP from one GPU, through a laptop-hosted server, to the next GPU shard.
  • Built an evaluation platform that breaks every request into end-to-end latency, per-shard prefill and decode time, vector transport overhead, and server-queue wait time.
  • Now optimizing a small model toward 15-20 tokens per second before testing a larger model across more distributed GPU networks.

Why I built it

EveryGPU began with a simple question: if a laptop, Colab, and Kaggle can each offer usable GPU capacity, why can’t they work together to run a model that none could serve alone? I started by exploring how scattered, otherwise-idle GPUs could act as one inference system.

The hard part

The first prototype made clear that adding GPUs is not enough. Each request must be split, scheduled, and carried between machines without letting queueing or network transfer erase the gains from extra compute. Right now, activation vectors travel over TCP through a laptop-hosted server between GPU shards.

What I learned

Distributed inference needs measurement before optimization. The evaluation platform separates queue wait, per-shard prefill and decode, and transport overhead, so I can see where a request actually spends time.

Where it goes next

The immediate target is an efficient small-model system at roughly 15-20 tokens per second. From there, I want to test direct peer-to-peer transport, potentially with QUIC, and scale to larger models and more remote GPUs. The longer-term work is routing requests well across n GPUs and m shards, sustaining useful concurrent throughput, recovering from unreliable nodes, and building a desktop app that gives a participant’s GPU back when they need it.

How it works

  1. Prompts

    requests enter the system

  2. Router

    splits and schedules work; a laptop server today

    activation vectors over TCP
  3. GPU shards

    spare laptop, Colab, and Kaggle GPUs, each running one model stage

  4. Tokens

    generated output

An evaluator records server-queue wait, per-shard prefill and decode, transport overhead, and end-to-end latency. Next: direct peer-to-peer transport, n GPUs by m shards, and recovery from unreliable nodes.

Next projectDictate