On the same engine, ten decisions read as single tokens take about 92 ms. Generating them as JSON takes about 1.57 s, and the cascade also spends about 411 ms on transcription.
DuplexJev
DuplexJev turns a frozen speech LLM into a decision engine for full-duplex voice agents. Turn-taking, which filler to play, and who is speaking are each read as a single-token distribution over option letters. Nothing is decoded. The product name is Speech-to-Decision.
The paper is submitted to IEEE ICASSP 2027. Acceptance has not been announced. This preprint link will move to arXiv when that version is up.
Decisions without decoding
With the cross-attention connector, gender and emotion both reach about 90%. Transcript distillation alone leaves them near 55% and 28%.
Spoken QA with the last-layer connector is 90%, against 91% when the model reads the transcript. The cross-attention connector drops spoken QA by one point (83% to 82%).
One forward pass for every question
- 01 Audio One utterance, or a batch of calls.
- 02 Frozen ASR encoder An off-the-shelf encoder such as Qwen3-ASR. Weights stay frozen.
- 03 Connector The only trained piece. Last layer, or cross-attention over three layers.
- 04 Frozen LLM One forward pass. Shared prefix, zero decode steps.
- 05 Single-token options Each question is a distribution over its letters. The max probability is the confidence.
Any ASR encoder and any frozen LLM join through a small connector. Supervision is on the answer token that is read out, while content keeps transcript distillation, so the model can hear what the transcript never wrote down.
Code, weights, and the product demo
The code, training recipe, batched inference pipeline, and the bilingual spoken-QA set qa100 are open. The deployable complete model is DuplexJev-4B.
The full-duplex demo is the product experience. It is not an online reproduction of this paper.