From Broken Subtitles to a Universal Video Script Platform

My Hackathon Journey at the Lingo.dev Hackathon
Have you ever downloaded a movie, only to find it’s in a language you don’t understand? You sit down to watch, excited—but there are no subtitles in your language. Suddenly, the whole movie is ruined, and you can’t follow a single scene.
That frustration is exactly what inspired my hackathon project.
The Mission: Why “Script-First” Matters 💡
While dealing with subtitles, I noticed a bigger problem:
Subtitles are usually created separately from the original video.
Because of this, translations lose context, timing breaks, and meaning changes.
So instead of fixing subtitles, I asked a different question:
What if we start with the script instead of the video?
Not random subtitle files.
Not direct video translation.
But a clean, accurate script generated from the original content, which can then be translated into any language.
That idea became the foundation of the Universal Video Script Generator & Localization Platform.
What the Platform Does (In Simple Terms)
The platform follows one clear pipeline:
Video → Script
Script → Translated Script
Script → Optional Subtitle File (SRT / VTT)
By focusing on text first, we remove dependency on unreliable subtitle sources and give users full control.
⚙️ Technical Decision Early On: Audio-to-Text Trade-Off
One of the most important engineering choices in this project was how speech-to-text is handled.
I used a local Whisper model via @xenova/transformers.
Why local Whisper?
Pros
Works offline
No audio sent to external servers
Full control over data and processing
Trade-Off
- Audio-to-text conversion takes more time for long files
This wasn’t a mistake—it was a conscious hackathon decision.
Accuracy, privacy, and control mattered more than raw speed.
Future versions can improve performance using:
GPU acceleration
Chunking and parallel processing
Optional cloud-based ASR
How It Works: Small Videos vs Large Movies
Video sizes vary a lot, so the system handles them differently.
🎬 Large Movies
Users upload existing subtitle or script files (SRT, VTT, TXT)
No need to upload multi-GB movie files
Faster, cheaper, and more reliable
🎥 Small Videos (Clips, Lectures, Interviews)
Users upload the video directly
Audio is extracted
Script is generated automatically using Whisper
This separation keeps the platform lightweight and scalable.
The Old Way vs The Script-First Way
| Problem | The Old Way | The Universal Script Way |
| Missing subtitles | Searching random websites | Auto-generate from audio |
| Large movie files | High bandwidth & storage | Text-only processing |
| Translation quality | Direct video translation | Script cleaning + Lingo.dev |
| Reusability | One subtitle per language | One script, many languages |
🧠 Tech Stack: A Deep Dive (System Design View)
🧠 The Brain
Whisper (local) via
@xenova/transformers– speech-to-textLingo.dev – accurate script translation
🏗️ The Backbone
Node.js – backend logic
Supabase – database & storage (text only)
🖥️ The Interface
Next.js + React – frontend
Tailwind CSS – clean UI
🎧 Media Processing
- FFmpeg / Web Audio API – audio extraction
Who Is This For? (Real Impact)
This platform helps real users:
🎬 Movie & TV viewers
Accurate subtitles in their own language🎓 Students & learners
Read and understand educational content easily🎧 Podcast & interview listeners
Turn spoken content into readable text♿ Accessibility users
Better access for hearing-impaired users
The Result: From Idea to Working System
The final system:
Stores only text, never video
Works with any video source
Generates reusable scripts
Supports multiple languages from one script
📊 Sequence Diagram

