eight ways of reading a video
It downloads the clip, reads every word on screen, transcribes the audio, checks the claims against the web, and measures how it did against everything else the creator has posted. Two to four minutes.
starterstory · 33.2s
5 live web searches ran for these, and every source below was fetched to confirm it loads. Searched for: Convex backend database pricing free tier, Claude Code Max plan $200/month 20x usage, Resend transactional email API company, Axiom.co logs observability platform pricing, fal.ai image video generation API pricing.
“We use Convex for our back end and database”
supported Convex is a real backend-as-a-service platform offering database, auth, and serverless functions, with a free starter tier and paid plans around $25-40 per user per month.
“We use Claude code on the 20x Max plan”
supported Anthropic's Claude Max plan does have a 20x tier priced at $200/month, offering roughly 20 times the usage limits of the base Pro plan.
https://intuitionlabs.ai/articles/claude-max-plan-pricing-usage-limits
“We use Resend for all our transactional and mailing list emails”
supported Resend is a real developer-focused transactional email API company founded in 2023 and backed by Y Combinator.
“We use Axiom for logs and observability”
supported Axiom.co is a real logs, metrics, and observability platform with a free personal tier and a paid Cloud plan starting at $25/month.
“We use FoulAI (fal.ai) for all our image and video generation”
supported Fal.ai is a real serverless API platform offering over 1,000 image, video, and audio generation models billed on usage.
This is a short promotional video from starterstory in which a founder of a nearly million-dollar SaaS business lists out the tools that make up his company's tech stack.
The voice walks through each tool in the stack in order — Convex, Vercel, Railway, Clerk, Resend, Axiom, OpenAI and Claude, fal.ai, and Claude Code — narrating what each one is used for. The screen echoes almost every one of those names and functions as on-screen labels, but it also adds a layer the voice never touches: a monthly price and a URL next to most tools, for example $40/mo next to convex.dev, $150/mo near railway.com, $25/mo by clerk.com, $100/mo by resend.com, $25/mo by axiom.co, $5k near fal.ai, and $200/mo appearing at 30.0s. None of these figures are ever spoken aloud.
A listener who only listened would never learn the specific monthly costs or the $5k figure tied to the image-and-video generation section that the screen displays alongside the tool names.
$5k is the only value on screen without “/mo” on it. If it is also /mo, the total is $5,540 and $5k alone is 90% of it. If not, $540/mo.

the frame a second in
“What's the tech stack behind this nearly million dollar SaaS in 2026?”
There is sound here, and someone talking over it
aac, 2 channels at 44100 Hz, averaging -27.1 dBFS with peaks at -9.1.
Cut timing not measurable: there is speech here, so the beats are not reliable enough to measure cuts against
69.5k views · 23.7× this creator's median
Posted 3d ago, which is about 23.1k views a day since it went up. Their median video does 2.9k.
Band: big. Green means at least three times the median.
the account is at 2.8×
Their most recent half of videos has a median of 3.7k against 1.3k for everything before it, across 459 videos and 1426 days.
459 videos over 1426 days. Median 2.9k views. Green is three times that or better.
Showing the newest 120 of 459. The figures above are computed over all of them.
33.2s · 192 kbps · 44100 Hz · ID3 v2.3 · cover from the platform's own thumbnail. These were read back off the finished file, not copied from what was asked for.
Gaurav launched Fastlane just two months ago. It’s already doing $69K...
starterstory
TikTok
Starter Story
2026
original sound
https://www.tiktok.com/@starterstory/video/7683952455289146637?_r=1&_t=ZP-99gHmjnL3mh
123 KB embedded
Everything below this line is reasoning on top of those numbers rather than something read off the file. It is a guess with its working shown, and it is labelled that way on purpose.
The whole clip is one shot (0 hard cuts, average shot length equals the full 33.2s duration), so this was most likely filmed as a single continuous vertical take on a phone rather than assembled from multiple clips in a multi-cut edit.
Only 6 of the 51 on-screen text regions move with the footage, meaning the text is mostly static overlay rather than content scrolling under a moving camera, which makes a screen-recording origin (e.g. a captured browser or app session) unlikely.
Speech fills 0.99 of the runtime with a longest inter-word gap of just 0.17s and a fast, even 207 words per minute, a cadence tighter and more metronomic than typical live-mic speech, which points toward a synthetic TTS voice track (tools like ElevenLabs) rather than a recorded human read.
27 of the 51 text regions echo spoken words, confirming burned-in captions generated from the audio and rendered into the frame, the standard output of an auto-caption feature like CapCut's word/phrase-pop caption presets.
Loudness range is narrow at 3.2 LU with a true peak of -9.1 dBTP, consistent with a single dry voice track and no music bed layered under it (a music mix would typically push the true peak closer to the -1 dBTP ceiling).
The file is only 0.38 MB at 92 kbps overall for a 720x1280 HEVC/AAC export, indicating aggressive compression for a small file size, i.e. an app-side re-encode rather than a high-bitrate export master.
4 live web searches ran for this.