The subtitling bottleneck nobody budgets for
If you run social for a brand that publishes in more than one language — or one that just has messy, code-switching, real-world audio — you already know the dirty secret of the creator stack: captioning is where the workflow dies. Not the filming, not the edit, not the thumbnail. The moment a client speaks Cantonese and English in the same sentence, or a podcast guest drops into Georgian, or a webinar audience is half in Shenzhen and half in Hong Kong, the tools most of us pay for quietly give up. They hallucinate names, flatten mixed speech into one language, or just refuse. That’s the gap Subanana is aiming at, and it’s worth a serious look for anyone whose content crosses a language border.
What Subanana actually is, and why it’s not another Whisper wrapper
The pitch, per the maker’s own forum post, is an AI subtitling and transcription platform built for “the way people actually speak” — mixed languages, code-switching, multiple speakers, mid-sentence switches. You paste a public YouTube link or upload a file (up to 30 GB on every plan, according to the maker’s reply to a commenter), and you get subtitles with a real editor to polish them, exported as SRT, VTT, TXT, or DOCX. There’s also a meeting-recorder mode that returns a transcript, speaker labels, and a summary, plus bilingual subtitle translation with glossaries for terminology consistency.
That’s the feature list. The interesting part is the architecture underneath it, and it’s the thing I’d tell any operator to read closely before comparing it to Descript, Otter.ai, or Happy Scribe.
The routing decision is the whole product
Kai, one of the builders, explained in the launch thread that the team routes each language to a different speech model rather than locking into a single vendor: “Great English accuracy often turns into mangled Cantonese or messy code-switching. So Live Captions uses the same routing as our subtitle tool: each language goes to the model that handles it best, not whichever vendor we happen to be locked into.”
That sounds like engineering trivia. It isn’t. It’s the reason this product exists at all. Every general-purpose transcription tool I’ve tested — and I’ve burned real hours in CapCut’s auto-captions, Descript, and Riverside’s transcription — is optimized for English first and treats everything else as a fallback. When you’re captioning a bilingual interview for Instagram Reels or TikTok, that fallback is exactly where your credibility dies. A wrong name in a caption is worse than no caption.
The maker, Aric Fung, framed the origin story plainly: the team is based in Hong Kong, “a city where a room is rarely all one language.” The usual fix at events is an interpreter booth and a stack of receivers — a budget most organizers don’t have. That’s the real market here, and it’s a market the big transcription players have largely ignored because it doesn’t scale as cleanly as English.
Where this fits in a social media operator’s actual workflow
Let me be concrete about who should care, because “supports more languages” is not, on its own, a reason to switch tools.
The repurposing pipeline is where the pain lives
If you’re running a content repurposing operation — one long recording chopped into vertical clips for TikTok, YouTube Shorts, and Reels — the transcript is the spine of the whole thing. It’s what you use to find the hook, write the on-screen text, generate the LinkedIn carousel, and produce the blog post. When the transcript is wrong because the speaker switched languages mid-thought, every downstream asset inherits the error. I’ve watched a single mis-transcribed brand name propagate into six pieces of content before anyone caught it.
Subanana’s glossary feature — the maker describes “bilingual subtitles with glossaries that keep your terminology consistent” — is the part I’d test first. Terminology consistency is the unglamorous thing that separates a professional captioning workflow from a hobbyist one. If your client’s product name is spelled three different ways across your captions, that’s a brand problem, not a subtitle problem.
Why TikTok and Reels creators should care more than LinkedIn ones
Here’s my take, and I’ll flag it as opinion: the platform where this matters most is TikTok, and it’s not close. TikTok’s distribution is unusually caption-sensitive — the algorithm leans on on-screen text and audio transcription for content understanding, and a meaningful share of viewers watch with sound off. If your captions are wrong, you’re not just losing comprehension, you’re potentially feeding the recommendation system bad signal about what your video is even about. On LinkedIn, captions matter for accessibility and dwell time, but the feed is text-first and the stakes are lower. If you publish bilingual short-form, the captioning layer is closer to infrastructure than to polish.
The live captioning angle is a different product wearing the same name
The launch also covers live captions: a QR code on screen, audience members scan it, pick their language, and follow along on their phones — reading captions or hearing translation. Up to five languages at once, two on the big screen, a subtitle box over slides, or a transparent layer piped into OBS, vMix, or ProPresenter. Every plan, including free, includes a 5-minute live session.
Two deliberate design choices stand out, both explained by the team and both worth understanding before you evaluate it:
No meeting bot. It listens to your microphone or your computer’s audio rather than joining a Zoom or Teams call as a participant. That’s a meaningful difference from Otter.ai and Fireflies, which typically join as a bot. For a conference stage, a classroom, a church service, or a hybrid meeting where you don’t want a “Subanana Bot” appearing in the participant list, that’s an advantage. It also means it can’t capture remote participants’ audio the way a bot can — a real limitation, not a footnote.
Captions arrive a moment late, on purpose. Kai: “Raw real-time output flickers, rewrites itself, and breaks mid-sentence, which is tiring to read on a phone. We chose a short delay so we can clean up the text before anyone sees it.” Aric quantified it in a reply: roughly half a second per sentence for the first pass, around a second when the line is enhanced with surrounding context and the event glossary. That’s a defensible trade. I’d rather have a caption that’s a beat late and correct than an instant one that mangles the speaker’s name.
What I’d actually borrow from this launch, regardless of whether you buy
Even if you never sign up, there are operational lessons here that apply to any multilingual content workflow.
Route by task, not by habit. The team’s core insight — that no single model is best at everything — applies to your whole stack, not just transcription. Most social teams default to one tool for everything because switching costs feel high. The teams that produce the best multilingual content I’ve seen deliberately use different tools for different languages and accept the friction.
Treat terminology as a first-class asset. The glossary feature is a reminder that consistency is a content strategy, not a formatting detail. Build a shared glossary for every client or brand you manage, and feed it into every tool that supports one.
Design for the handoff, not the happy path. Louie Lee, an engineer on the team, described the scenario the product was built around: “the keynote ends, a guest speaker takes the stage in another language, and part of the audience suddenly needs a different way to follow along.” That’s the real-world moment most event and webinar workflows fail. The team’s answer is being able to change source and audience languages mid-session via “Edit event” — saving briefly pauses and reconnects captions, subsequent speech uses the new settings, and earlier caption history stays intact. If you run any live content, that’s worth stealing as a mental model: plan for the transition, not just the setup.
Where my judgment says it falls short
I want to be balanced here, because the launch thread is genuinely enthusiastic and the team is responsive, but enthusiasm isn’t a substitute for scrutiny.
Pricing is only partially disclosed. The maker says live captions are on the “Max plan” and offers a code, PRODUCTHUNT20, for 20% off the first three months, valid until 10 October. But the actual price of the Max plan, or any plan, isn’t stated in the source. If you’re evaluating this for a team, you’ll need to check the site directly. I can’t tell you whether it’s competitive with Happy Scribe or Rev because the numbers aren’t in the launch material.
The no-bot decision cuts both ways. It’s a feature for stage events and in-person rooms. It’s a limitation for remote-first teams who want a bot to join a Zoom call and capture everyone. If your workflow depends on that, this isn’t your tool — at least not for meetings.
The latency is real. A second of delay is fine for a keynote. It’s noticeable in a fast-paced Q&A or a rapid-fire panel. The team made a deliberate trade and I respect the reasoning, but “roughly a second” is a number you should test against your own use case before committing.
No reviews yet. The Product Hunt page shows “No reviews yet” — the enthusiasm in the thread is from commenters and beta users, not from a broad user base with track records. That’s normal for a fresh launch, but it means the reliability claims are unproven at scale.
The “who it’s NOT for” question. If you publish exclusively in English, this is probably overkill. CapCut’s auto-captions or Descript will do you fine, and you already know them. The value here is concentrated in multilingual, code-switching, and under-served-language content. Elene Kvitsiani’s comment about Georgian being “usually problematic” and Adana Marukhyan’s note about Armenian being a “constant problem” are the real signal — those are the users this was built for.
What I’d watch / test next
If you’re curious, here’s what I’d do this week, in order:
Run one real recording through the free tier. Every plan includes a 5-minute live session, so there’s no excuse not to test it on actual content rather than a demo. Use the messiest audio you have — a bilingual interview, a noisy panel, a code-switching client call.
Test the glossary feature specifically. Feed it three brand names and one technical term, then check whether they survive translation intact. That’s the feature that determines whether this is a toy or a tool.
Compare the transcript against your current tool on the same file. Don’t trust the launch copy; trust your own diff. If you use Otter.ai or Descript today, run the same recording through both and count the errors.
If you run events, test the QR-code flow with a real audience. The 5-minute free session is enough to see whether the phone-based caption experience holds up in a room with bad Wi-Fi.
Watch the pricing page and the review count. Both are thin right now. If the team publishes transparent pricing and real user reviews accumulate over the next quarter, that’s your signal to take it seriously for production work.
The broader takeaway: the creator economy’s multilingual moment is here, and most of our tooling still assumes English. Subanana is one of the first products I’ve seen that treats code-switching as the default rather than the edge case. Whether it wins or not, that’s the right bet — and it’s a reminder that the next wave of content tooling will be won by whoever handles the messy, real-world audio the incumbents gave up on.





