Sahaay

सहाय · offline lecture companion

Live captions, translation into 22 Indian languages and a jargon glossary for any lecture on your screen — all of it running on the machine, with the network switched off.

Built for the Snapdragon® AI Lab Build & Present Challenge 2026. The browser version downloads Whisper once and runs it on your own machine — no audio leaves the tab, which is the same claim the desktop app makes. Pick a language and it translates too, into any of 22 Indian languages, on your machine. The NPU and system-audio capture need the desktop build.

Sahaay · recorded session · Hindi
The Sahaay app mid-lecture: English captions, each translated into Hindi with technical terms such as eigenvalue kept in English, and a jargon panel explaining eigenvector and eigenvalue.
हिन्दी translated on the device eigenvalue stays eigenvalue 0 bytes of audio leave it
The real desktop interface, replaying a session it recorded.
1.55real-time factor on a CPU with the glossary running — the captions fall behind 13.47 msWhisper encoder on the Snapdragon X2 Elite NPU 129 / 129layers on the Hexagon NPU, zero CPU fallback 0outbound connections in a full session — tested, not asserted
THE FILM · 2:00

One lecture line, followed through the product

Captioned as it is spoken, translated into all 22 languages with eigenvalues kept intact, the network switched off mid-lecture, and the measurement that makes this a Snapdragon application. Every caption and translation in it is from a real recorded session. The score is original; the lecturer and the narrator are synthesised voices (Kokoro text-to-speech), and the player has captions.

Full quality (1080p, 60 fps) on GitHub: release submission-v2. Made frame by frame from the project's own UI and data: how.

01

The problem

A student in a regional-language school sits in an online lecture delivered in fast English, dense with terms that have no everyday translation — eigenvalue, determinant, characteristic equation. Captions help. Captions in their own language help more. Captions that also explain the jargon are the difference between following the lecture and losing the thread.

Cloud captioning can do this. It also requires bandwidth the student may not have, and it means every lecture they attend leaves their machine. Sahaay does it locally instead: three models, one laptop, no network.

02

Try it without installing anything

The browser build downloads Whisper once and runs it on your machine, in the tab. Your audio never leaves it.

That is not a compromise on the argument — it is the argument, on hardware you already have. The page is static files; the inference happens on your own GPU via WebGPU, or your own CPU via WebAssembly, and the badge in the header names which. Speak into your microphone, or share a browser tab playing a lecture and it will caption that instead.

Pick a language and the captions are translated as well — into any of 22 Indian languages: Hindi, Tamil, Telugu, Kannada, Malayalam, Bengali, Marathi, Gujarati, Punjabi, Odia, Urdu, Assamese, Maithili, Nepali, Kashmiri, Sindhi, Manipuri, Bhojpuri, Awadhi, Chhattisgarhi, Magahi and Mizo. It is the model the desktop app ships, NLLB-200, running in a worker on your machine, with technical terms protected so "eigenvalue" stays eigenvalue. It is about 900 MB, so it downloads only when you pick a language, and once.

Jargon is explained by the same seeded vocabulary the desktop app falls back to when no language model is installed — exported from the same source file, so the two cannot drift.

Watch the RTF badge while it runs. It should hold around 0.5: on the laptop this was tuned on, WebAssembly transcribes the whole test lecture at 0.46–0.50, so the captions keep up. That is one model on the CPU. The measurement below is what happens when a second model, the glossary, runs on the same cores — the part a browser cannot show, and the reason this belongs on a Snapdragon PC.

What the browser build is not. It uses Whisper Base rather than Small, and captures one shared tab rather than all system audio. Translation shares your CPU with Whisper, so each line's translation lands about six seconds after its caption (the captions themselves still keep up), and the 900 MB download is not one for a phone connection. And there is no Hexagon NPU in a browser, so it can show you the problem but not the fix. Those need the desktop build.

03

Why this needs an NPU

Turn the glossary on, and a CPU stops keeping up with the lecturer.

This is the measurement the project is built around. Sahaay runs Whisper and a language model at the same time, for an hour. On one set of CPU cores they compete directly, and the captions — the thing the user is actually reading — are what lose.

Whisper Small on real recorded speech, with and without the glossary model running. CPU, x86.
Whisper latencyMeanp95Real-time factor
Glossary idle3275.9 ms3422.3 ms0.409
Glossary generating12400.2 ms13241.6 ms1.55

1.0 is the line that matters. Below it, transcription finishes faster than speech arrives. Above it, every minute of lecture takes more than a minute to caption, so the captions fall further behind and never recover. Moving the Whisper encoder onto the Hexagon NPU removes exactly this contention, and it is why the application belongs on a Snapdragon PC rather than anywhere else.

This number was wrong until 22 Sep 2026. The harness reported 0.043 → 0.299 and concluded the pipeline still kept up. It had measured whisper_tiny_en against a synthetic test tone and recorded neither fact — the model id is auto, so what gets benchmarked depends on which weights happen to be on the machine. The product resolves to whisper_small_portable. Both harnesses now print the model and the signal beside every number.

Method and full results →

04

Measured on Snapdragon silicon

Sahaay was developed without a Snapdragon PC on the desk. Rather than claim NPU performance that could not be verified, the graphs went to Qualcomm's own device farm. Every number has its AI Hub job ID — the job pages open after signing in with a Qualcomm ID, and the full profiles are in docs/AIHUB.md.

Whisper encoder, one full 30-second mel window, via Qualcomm AI Hub.
DeviceInferencePeak memoryLayers on NPUJob
Snapdragon X2 Elite CRD13.47 ms9.2 MB129 / 129 jglyo70e5
Snapdragon X Elite CRD27.56 ms17.0 MB129 / 129 j5m0o4zyg
Snapdragon X Plus 8-Core CRD26.76 ms16.8 MB129 / 129 jpxl3m7jp

129 of 129 layers execute on the NPU — nothing falls back to the CPU. The decoders are deliberately absent from this table; they carry a dynamic KV cache, so a single fixed-shape profile would measure one arbitrary sequence length and read as though it were the cost per caption. It is not.

All device-farm results →

05

How it fits together

CPUWASAPI loopbackCaptures whatever is playing — Zoom, YouTube, a recording. No virtual cable, no driver install.
CPUSilero VADSplits on pauses, not on a fixed clock, so captions break where sentences do. Tiny, and on the CPU by design.
NPUWhisper SmallTranscription and per-segment language detection. Its encoder is the graph measured above — 129 of 129 layers on the NPU.
NPU targetNLLB-200Translation into 22 Indian languages, with technical terms protected from being translated into nonsense. Runs through the QNN provider where it is available; not separately profiled.
NPU targetLlama 3.2Jargon glossary while the lecture runs; notes and a self-test when it ends. Ships Hexagon assets for Snapdragon; falls back to an int4 CPU build.

NPU measured on Snapdragon silicon   NPU target built for it, not yet profiled   CPU stays on the CPU by design

Architecture in detail →

06

The detail that matters most

A student can work around a garbled function word. A garbled eigenvalue breaks the link to the textbook — so technical terms are scored separately from everything else, and protected through translation rather than left to chance.

SetUtterancesWord error rateTechnical terms kept
English40.0%8 / 8
Code-mixed363.6%6 / 6
All723.7%14 / 14

Read this honestly: the audio is synthesised by Windows SAPI, not spoken by people, and only an en-US voice is installed — so the code-mixed row is read with English phonetics. These numbers are a floor, not a forecast. Expect materially worse on a real lecture recording.

Per-utterance results →

07

Does it survive a whole lecture?

+0.01 MBmemory growth per simulated minute, over 60
4 → 4threads, start to finish — nothing leaks
0segment-queue backlog at the end
400+tests, green on x86, ARM64 and Linux

Every other test in the repo runs for seconds. The failure modes of an hour are different in kind: unbounded lists, queues that only grow, event history that accumulates, threads that leak. So the soak test reports slope, not a single number.

Soak method →

08

Offline is tested, not asserted

"Runs offline" is easy to put on a slide. In this repo it is a test: tests/test_offline.py patches the socket layer so any non-loopback connection or DNS lookup raises, then runs a full session and asserts nothing tried. The guard self-tests first — a guard that never fires would make the whole test vacuous.

09

Run it yourself

Windows on ARM64 or x86. The installer picks the right model tier for the machine it lands on.

git clone https://github.com/andringodson/Hackathon-SnapdragonAILab
cd Hackathon-SnapdragonAILab
.\install.ps1
.\run.bat

Then .\run.bat --selftest reports what works, what is degraded and what failed — and every line that is not OK says what to do about it. No Snapdragon device? It runs on CPU, more slowly, and the badge in the header says so rather than pretending.

10

What is not proven yet

Stated here rather than left for a judge to find:

  • The full three-model pipeline has never run on physical Snapdragon hardware. Individual graphs have, on Qualcomm's device farm, at 129/129 layers on the NPU. The end-to-end run on one machine has not. The suite does now run on ARM64 Windows in CI, which is the same instruction set but Microsoft silicon with no Hexagon — it proves the install path, not the NPU.
  • Word error rate is measured against synthesised speech, not real speakers — no accent, no room, no crosstalk, no disfluency.
  • The NPU column of the concurrency table is empty. The measurement above is the CPU half of the comparison; the half that closes the argument needs a Snapdragon PC and one command.