Skip to content
DUETA
Blog

Notes from the pipeline

How DUETA takes a recording of two people talking over each other and returns one clean track per speaker. The models, the ordering, the measurements, and the parts we can't yet put a number on.

Latest6 min readPipelineSpeech separation

Separating the Inseparable

What actually happens to a two-speaker recording between the upload and the two tracks that come back.

Read
Archive
20 August 2026

Scoring separation without ground truth

SI-SDR needs a clean reference recording. In production there is never one. Here is the model we use instead, and how far to trust it.

7 min readEvaluationSI-SDR
19 August 2026

One endpoint, five models

We shipped the pipeline as five callable stages, watched people use it, and collapsed it into one job. Here is what that cost and what it bought.

7 min readInfrastructureAPI design