New Delhi: Gnani.ai which is an Indian AI startup, they have recently introduced its latest five billion parameter voice to voice foundational AI model at the Indian AI Impact Summit 2026. And this latest model, called Inya VoiceOS, has been built under the IndiaAI Mission, and it is being released as a research preview ahead of the larger 14 billion parameter voice-to-voice model that Gnani.ai is developing. Inya VoiceOS, which also supports more than 15 Indian languages, produces 24 kHz audio output, and it is especially designed to operate with sub-second end-to-end latency.
Inya VoiceOS is especially designed to operate directly on the speech instead of relying on the more common pipeline of converting speech to text and then back to speech. Gnani.ai stated that the model works in what it describes as acoustic and semantic space, by enabling it to consume and generate speech tokens without intermediate transcription steps. Model jointly encoded phonetics, prosody, semantics, and intent, and is built to retain paralinguistic cues such as tone, emotion, pacing, and pauses. It also supports streaming and interruption-aware inference, and can handle overlapping speech and mid-utterance corrections without resetting the conversation.
Gnani.ai is stated to have Inya VoiceOS, which has 5 billion parameters and has been trained on more than 14 million hours of multilingual speech data, with an additional 1.2 million hours of task-specific fine-tuning data. The model has also been trained using over 8 trillion text tokens for linguistic grounding and reasoning. It supports more than 15 Indian languages, produces 24 kHz audio output, and is designed to operate with sub-second end-to-end latency, including in code-mixed speech scenarios.









