Speech translation has matured impressively. Multilingual streaming models report strong results on specific language pairs and benchmarks (SEAMLESS, Nature 2025). But the progress is lopsided, and the lopsidedness turns into a problem everyone feels and few name: the listening side keeps getting freer while the speaking side stays chained to a device.
Understanding a foreign language is now the comfortable half. The phone sits on the table, speech is recognized, the translation appears on screen or in an earbud. Then it is your turn to answer. You lean into the phone, speak, wait, and play the result to the other person. If something comes out wrong, the loop starts over. This loop costs three different things.
It costs identity. The voice that reaches the other side is the device's, not yours. A human voice carries more than words — age, emotion, resolve, warmth, all coded in timbre and rhythm. Today's translation strips that layer off. The listener hears what you said and never hears you.
It costs time and rhythm. Turn-taking in live conversation is startlingly fast; cross-linguistic research puts typical gaps at the scale of a few hundred milliseconds (Stivers et al., PNAS). Translation carries an unavoidable structural delay — languages order words differently, which is why even simultaneous interpreters run a few words behind and lean on anticipation. Stack the speak-wait-play interface loop on top of that, and the natural pulse of dialogue dies.
And it costs the body. A conversation held with one hand on a phone and both eyes on a screen loses what makes conversation physical: eye contact, free hands, the ability to interrupt. A tourist cannot haggle like that. A migrant cannot plead a case. A nurse cannot translate while holding a patient's hand.
What needs solving is not the translation of words — that front is strong and moving. It is the act of speaking itself: letting it pass through translation while keeping its owner's identity, rhythm and physical freedom. That is Omniloq. And a description of Omniloq must start with an honest label: this is not a working prototype. It is a research program whose feasibility has been examined systematically, whose limits have been drawn, and whose first experiment is defined.
The framework came out of a multi-perspective feasibility review — physics, signal processing, neuroscience, speech technology — in which independent assessments were collected and their disagreements argued through. The most valuable output of that process was a clear picture of what cannot be done, because workable designs begin where the limits are honest.
First limit: sound, once radiated, cannot be recalled. What leaves a speaker's mouth spreads through the room; erasing it for every listener, or rewriting its content mid-air, is not a practical option (interference-based cancellation works in confined zones like an ear canal, not in open space). The design consequence is severe. If the user speaks at normal volume, the other person hears the original language first and the translation a second or two later, stacked on top of each other — and the promise of natural conversation collapses. Silent or near-silent articulation is therefore not a feature of Omniloq. It is the mandatory input. The user articulates below a whisper, almost without audible sound; nothing meaningful leaks into the room; only the translated voice reaches the other side.
Second limit: translation needs the future. A sentence's translation often cannot be fixed until the sentence ends. The delay cannot be eliminated; streaming systems show it can be pushed down to around two seconds with anticipation (Meta Seamless), a figure that varies by language pair and conditions. Spoken output adds a harder twist: it cannot be un-said. Text interfaces revise the previous word silently; audio corrections are heard. So prediction alone is not the answer. Commit policies — when the system decides a phrase is stable enough to speak — and an audible repair protocol for when it commits wrongly are first-class design and measurement targets, on par with average latency.
Within those limits, the architecture has three layers.
Capture. The speech chain leaves traces that do not travel through the air, and skin-surface sensors can read them: tissue vibration at the throat (physically still an acoustic signal — structure-borne, conducted by contact rather than airborne) and the electrical activity of articulation muscles, surface EMG. These channels resist ambient noise well — not perfectly; motion artifacts, clothing friction and placement drift are real — and controlled experiments show they carry information even in silent articulation (Gaddy & Klein, EMNLP 2020). Honesty is required here too. Contact signals are narrow-band and miss parts of certain phonemes (TAPS study); the system does not "read" speech so much as reconstruct it, filling gaps with a language model. The unavoidable risk is semantic hallucination, and it is managed with targeted metrics rather than headline accuracy: error rates on numbers and proper names, negation reversal, the frequency of words the user never said. In high-stakes settings — health, legal, official — the acceptance thresholds must be separate and far stricter. (Neural capture, decoding speech intent through implants, has been demonstrated experimentally (Nature Human Behaviour), but it belongs to invasive research with clinical populations. For Omniloq's consumer scope it stays where it is: a long-horizon scientific frontier, out of scope.)
Decision. The system must voice what the user chose to say — never a murmur, an inner rehearsal, half a sentence. This does not have to be solved by AI alone, and should not be. Physical and semi-physical activation — touch to engage, a deliberate jaw gesture, a hard mute switch, two-step intent confirmation — cuts false triggering at the architectural level; automatic intent detection sits on top of those guarantees, not in their place. The metric is blunt: false activations per hour, measured and under threshold before the field.
Production. The translated content is re-synthesized in the user's own voice, from a profile recorded in advance with explicit consent, and played from a small speaker on the user or the listener's earbud. Two honest notes. Identity-preserving synthesis is not in itself new — systems preserving timbre and emotion through translation exist; if Omniloq has an originality claim it lies in the end-to-end integration with silent-articulation input, and even that claim waits on a systematic literature and patent search before anyone states it as a first. And a voice profile is biometric data with obvious abuse potential. On-device storage and encryption, no profile export, refusal to voice any text the user did not produce, watermarking of synthetic output, and the right to delete the profile are the non-negotiable security floor. Directed-audio tricks ("only the listener hears it") may someday be a comfort feature; reflections and alignment problems mean they can never be sold as a privacy guarantee.
A full two-way conversation adds its own engineering layer — capturing the other side's reply, separating languages and speakers, managing turn-taking, cancelling echo. The noise-robust listening core for that layer comes from Omniloq's sibling under the Sonverto roof: Sonacern's reference-conditioned extraction engine is the listening half. The two projects are one vision from two directions — hear in crowds, speak at the source.
The program's first experiment is defined and deliberately modest: very quiet speech, a throat-contact microphone, and existing translation and synthesis chains. No neural layer, no custom hardware. The experiment carries a known tension inside it, and measuring that tension is largely the point: as voice level drops, the throat signal weakens too — and "very quiet" is not guaranteed to be inaudible at close range. The trial will put numbers on the trade-off: at what level does the signal remain decodable, and what does a listener one meter away actually hear? If low-voice mode proves insufficient, capture weight shifts to the sEMG layer — already defined as plan B. Whatever the results are, they will be reported with their context, the same way Sonacern's were.
On human dignity. The language barrier is among the most widespread and least visible inequalities of the modern world. A person forced to speak through a third voice — an interpreter, their own child, a device — surrenders a piece of their expression every time. Describing your symptoms in a hospital, your situation at a government desk, your joke at a market stall, in your own voice: that is not comfort, it is the integrity of a person's social presence. When speech can change language while keeping its identity, translation stops being a service someone provides you and becomes something closer to your own ability.
On the integrity of communication. Research on conversation keeps finding that dialogue owes much of its efficiency to channels beside the words — millisecond-scale turn-taking, timbre, gaze. Screen-mediated translation cuts all of them; a speaking-side solution aims to give exactly those back. The claim of "natural conversation" should live or die by measurement: the turn-gap difference between device-mediated and unmediated dialogue, balance of speaking time, intelligibility, and the safety and accuracy thresholds above. The day those numbers are met, the claim stops being rhetoric.
On the science. The speaking-side problem sits at the crossing of research lines that matured separately — contact sensing, silent-speech interfaces, streaming translation, identity-preserving synthesis. An integrated system forces questions none of the parts face alone. How is a latency budget split across layers? How is a prediction error repaired in already-spoken audio? At what guarantee threshold may intent detection go to the field? In which content classes is gap-filling from an incomplete signal never acceptable? Measured answers to these questions travel well beyond this project — toward speech prostheses for people who have lost their voice, and the future of accessibility at large.
Physics fixes two limits: radiated sound cannot be recalled, and translation needs the future. Both are permanent. But the same physics has already shown — in separate, measured experiments — that the non-airborne traces of speech can be read, that content can be translated near real time, and that identity can be preserved in synthesis. Omniloq does not ask for a miracle beyond those parts. It asks whether they can be made one honest, measured whole. People talking to each other with no screen between them, in their own voices, whatever the language: that is the goal, and every step toward it will be taken the way this document is written — with the proof attached.