Skip to content
Home » Blog » AI News — August 12, 2026: DeepMind ships sign language translation to phones

AI News — August 12, 2026: DeepMind ships sign language translation to phones

The big story: SL2T puts sign language AI in users’ hands

Google DeepMind announced SL2T today, a massively multilingual sign-language-to-text translation model — and, more notably, shipped it. SL2T now powers sign-to-text dictation in Gboard and Live Transcribe on the Pixel 11, starting with American Sign Language to English. It’s the first time sign language AI has made it out of the research lab and into a consumer product at this scale.

The practical framing is simple: Deaf users can now sign to their phone anywhere they’d otherwise type. Search the web, draft a message, prompt Gemini, or respond in a Live Transcribe conversation without typing back and forth. DeepMind says testers found signing in ASL faster and more natural than typing in English. The feature costs nothing extra, and more devices are promised.

Why this was hard

DeepMind is blunt about why progress here has lagged spoken-language AI by decades. Speech transcription is a sequential mapping from sound to text within one language. Sign languages aren’t that — there are more than 200 of them worldwide, used by an estimated 70 million Deaf and hard of hearing people, and each is an independent natural language with its own grammar and lexicon. Translating ASL to English is genuine machine translation, not transcription.

The second problem is perception. Meaning is carried by simultaneous movement of hands, arms, torso, head, and face, tracked at high frame rates. That’s an expensive computer vision problem. It’s also why earlier attempts like sign language gloves were doomed: sign language is not, as the post puts it, “English on the hands.”

How SL2T works

Three design choices stand out.

Scale across languages, not just one. The model trained on over 100,000 hours of data spanning more than 50 sign languages, with roughly a quarter of it ASL. Training jointly across languages, dialects, and proficiency levels let the model learn shared structure — and beat single-language models in DeepMind’s experiments.

Landmarks, not video. SL2T never sees the camera feed. An on-device model (MediaPipe Holistic) extracts pose landmark coordinates, and only those geometric points go to the server. The video is discarded immediately. Privacy architecture and efficiency in the same decision.

No glosses. Prior work leaned on “glosses,” intermediate word-level annotations that flatten out non-manual markers and spatial constructions. SL2T translates landmark sequences straight to text, which removes artificial vocabulary ceilings and lets quality scale with data.

On FLEURS-ASL (sd-test), SL2T scores 70 BLEURT zero-shot — well above any previously reported result. But the more telling work is the unglamorous part: minimizing streaming latency, preventing hallucination when nobody is signing, ensuring fairness for the ~10% of signers who are left-handed, and handling one-handed signing, since you’re holding the phone with the other hand.

The governance angle

The project was conceived by Sam Sepah, a Deaf Googler, and DeepMind stood up an AI Sign Language Advisory Committee (AISLAC) of global Deaf organizations and subject-matter experts to guide deployment. They co-published a joint impact report documenting capabilities and limitations, and say they’ll do the same for future releases. Worth watching whether that holds — participatory governance is easy to announce and hard to sustain.

Next up, per DeepMind: more sign languages, sign language generation, and frontier model capabilities.


Also this week

DeepMind’s leadership shakeup is still reverberating. Demis Hassabis moved to chairman of Google DeepMind and Alphabet chief scientist on August 5, with CTO Koray Kavukcuoglu taking operational control as SVP reporting to Sundar Pichai (CNBC, Time). Alphabet shares fell nearly 4%. Fortune’s follow-up reports Gemini 3.5 Pro has missed three release deadlines, five senior researchers left in a single week in June, and engineers are working 60-hour weeks — retention bonuses notwithstanding.

Jeff Dean and Sanjay Ghemawat left to start Discovery Loop. Google’s 30th employee and his longtime collaborator departed after 27 years, joined by Google Brain founding member Quoc Le and DeepMind’s Oriol Vinyals, to found a public benefit corporation aimed at automating complete experimental loops — running thousands of experiments in parallel, starting with ML research itself. Dean is CEO. Radical Ventures and Khosla Ventures are co-leading the round, with Kleiner Perkins, Lightspeed, Doerr Capital — and Alphabet itself — participating (TechCrunch, Jeff Dean’s announcement). The company is also openly interested in recursive self-improvement, which is either the point or the problem depending on who you ask.

WeatherNext gained a day of cyclone forecasting lead time — and got open-sourced. A Nature paper published August 6 shows WeatherNext Cyclones hitting state-of-the-art on track, intensity, and wind structure simultaneously — three-day forecasts as accurate as prior two-day ones, roughly a decade of meteorological progress. It runs 1,000-member ensembles and produces a 15-day forecast in under a minute on a single TPU. Code and weights are on GitHub.

Gemini Robotics 2 moved past the tabletop. Announced July 30, the release covers whole-body control, five-finger dexterity, and multi-robot collaboration across three models — the VLA, Gemini Robotics ER 2 for embodied reasoning, and an on-device variant. All three are in early access.


Sources are linked inline. Corrections welcome.

Tags:

Leave a Reply

Discover more from MONTANIMATION

Subscribe now to keep reading and get access to the full archive.

Continue reading