Nari Labs provides optimized inference for multimodal models, focused on speech workloads like real-time TTS and streaming STT. It’s aimed at teams that need predictable latency and lower serving costs without spending as much effort on infrastructure work.
If you’re building an app that depends on speech in real time, like low-latency voice experiences, Nari Labs is positioned as a way to get fast responses from TTS and STT while keeping production work manageable. The company also offers a pathway for customization through fine-tuning, rather than only using hosted endpoints.
Nari Labs also highlights in-house model work tied to its inference stack, including (open-source dialogue TTS) and (a human-centric avatar model with streaming and non-streaming variants). That pairing of model assets with dedicated inference is the main reason to consider it if you want both serving performance and options to tailor behavior to your domain.
+3 more
+2 more
+2 more