Say something.
Tap the button or press Space. Words appear while you are still talking and settle when you pause. Your audio is transcribed on the Jetson and goes nowhere else.
ssh -L 8770:localhost:8770 agx-ts and use localhost.
Tap the button or press Space. Words appear while you are still talking and settle when you pause. Your audio is transcribed on the Jetson and goes nowhere else.
The last few models you used stay loaded on the Orin, so switching between them is instant. The first pick of a new one downloads it once.
Each translation is spoken aloud in its own language, by the Jetson. Pick a voice for every language; choosing one plays a short sample.
The microphone pauses while the device speaks, so it never transcribes itself.
For the natural voices: they follow a description of how to speak.
Asking the Jetson which voices it has…
Trade a little patience for whole sentences, or the other way round.
Used when Language is "Auto · Indian + English". The fewer languages, the fewer mix-ups: tick only the ones your speakers use.
When someone speaks the language the page translates into (English, translating into English), their line is kept as it is by default: shown and spoken in that language. Or send it the other way: into the other language of the conversation (the one heard most lately, Hindi until another is heard), or into a language you pick.
How long a pause settles a line. Shorter settles sooner but can split sentences.
Re-transcribes back-to-back while you talk. Off saves GPU.
Top bar: the microphone, with a mark at the room's noise floor. Bottom bar: how sure it is that this is speech — a line starts when it crosses the mark.
How sure it must be before it starts a line. More eager catches a quiet talker; too eager and a fan becomes a sentence. Watch the bottom bar while you speak.
Lifts a quiet voice before anything listens to it. Measured: 24 dB takes a talker at −55 dBFS from 56% of their speech detected to 79%.
Sound this far below the person speaking counts as silence, so someone across the room stops holding your line open.
Mixes a denoised copy into what the recognisers hear, never the detector. Helps against a fan or nearby talk; costs a few points on clean speech.
Each line can be played from the transcript. Off plays the microphone as it arrived; on plays the same audio after the gain and the clean-up, which is what it actually worked from.