CafeFX.AI
A 3-tier speech-to-text fallback: why the Mac Mini became the fastest node in our audio pipeline

A 3-tier speech-to-text fallback: why the Mac Mini became the fastest node in our audio pipeline

Our daily short-video pipeline has one quiet but critical step — turning speech into text (ASR) before editing, subtitling and quality checks. This post documents a small change with a big effect: reordering our 3-tier ASR fallback so the fastest respondent runs first.

The problem: ASR is an invisible bottleneck

Whenever the primary transcriber is slow or its queue is long, the whole day's batch drags behind it. We used to put the render GPU first, assuming GPU always wins. The real numbers said otherwise — whenever the GPU was busy rendering video, ASR jobs waited far longer than necessary.

The new 3-tier architecture

The idea is simple: let the idle, fast machine answer first, then fall back layer by layer — automatically.

Tier Machine Role Timing
1 Mac Mini (MLX) Primary — fast, never fights the render GPU ~3.4s first token
2 GPU server First fallback timeout 90s
3 CPU on the main box Last fallback timeout 300s

The key is not just speed but workload separation — tier 1 draws from a different resource pool than rendering, so both queues move in parallel.

Measured results

After the swap (a single reorder inside one config file), the system reported:

Live-probe output of the 3-tier ASR system
Post-swap live probe: tier 1 answers in ~3.4s while both fallback tiers stay on standby
  • Text similarity versus the reference transcript held at 0.9956 — identical to before the reorder, so quality was not traded for speed
  • The filler-word gate passed across the full set
  • End-to-end checks came back green

Lessons from the reorder

Three shortcuts for other teams:

  1. Measure before you believe — "GPU is faster" is only true when the GPU is idle. Numbers from the live system beat assumptions.
  2. Separate workloads; don't let them fight — rendering and transcription should not share one queue.
  3. Layered fallbacks with explicit timeouts — clear cutoffs let the system self-heal without a human watching.
Render queue screen on the GPU that competes with transcription jobs
The render queue — the reason the new ASR tier lives on a separate machine instead of fighting this queue

The ASR tier sits behind our central model gateway, which handles queueing and access control for every AI service we run.

Admin screen of the team's LLM gateway
The admin screen of the central gateway that ties our models and AI services together

What's next

The next piece of this work is automatic audio quality gating before transcription (noise and level checks), so unusable jobs get rejected at the door instead of failing mid-queue. Results will be documented here.

แหล่งอ้างอิง

เริ่มต้นเทรดทองคำและ Forex กับ XM

เปิดบัญชีผ่านพาร์ทเนอร์ CafeFX (Partner Code: cafefx) รับการดูแลจากทีมที่มีประสบการณ์กับนักเทรดไทยมากว่า 13 ปี

เปิดบัญชี XM ฟรี →
แอป iCafeFX
iOS · App Store Android · Google Play

การลงทุนมีความเสี่ยง ผู้ลงทุนควรศึกษาข้อมูลก่อนตัดสินใจลงทุน

Part of the iCafeFX Network · iCafeForexXMSignalSiamCafeSiam2RSiamLanCardSiamCafeBookKittithatCafeFXRedhatAI