Global Tech News Technology news from original sources.
AI

RetroThinker Lifts Voice AI Math Accuracy by 11 Percentage Points

Digital audio recorder on a wooden table, used as an illustrative image for speech AI research

A training method called RetroThinker raised a voice AI model's accuracy on spoken mathematics problems by 11 percentage points with almost no change in response delay. University of Texas at Austin and Meta researchers compared it with the same model beginning its calculation only after the entire question had been heard. The work was posted as an unreviewed preprint on September 10.

The underlying Moshi model processes incoming speech, outgoing speech and an internal text stream at the same time. That design allows it to listen and speak simultaneously. RetroThinker uses the internal stream to begin a calculation while the question is still arriving instead of leaving every reasoning step until the speaker stops.

The model writes provisional calculations, checks them as more words arrive and corrects them when the developing question changes their meaning. Training first exposed it to deliberately altered calculations and then to mistakes the model made itself. A final stage rewarded correct but shorter solutions so that earlier reasoning did not simply turn into a longer pause before the spoken answer.

The evaluation used synthetic speech reading questions from GSM8K, a collection of school-level mathematics problems. Another AI transcribed and graded the spoken answers. Under those conditions, beginning the work early delivered the reported gain without materially increasing latency.

Natural voices, background noise and interruptions remain outside the evidence, and accuracy was weak on problems requiring six or more steps. Human-checked scoring and live conversations are needed to establish whether the method preserves its advantage when the speech itself is uncertain.

Sources