NOTICE — the demo audio in this directory The twelve clips are one 10.8-second moment of a REAL recorded meeting, taken from the NOTSOFAR-1 dataset and separated by this product's own public API: every one of the meeting's six microphones, before and after. SOURCE. Microsoft's NOTSOFAR-1 ("Natural Office Talkers in Settings Of Far- field Audio Recordings") dataset — real, unscripted meetings recorded in real conference rooms, released for the NOTSOFAR-1 / CHiME-8 speech research challenge: https://huggingface.co/datasets/microsoft/NOTSOFAR The exact material: meeting MTG_30880 of the dev set (version 240825.1_dev1, path benchmark-datasets/dev_set/240825.1_dev1/MTG/MTG_30880/close_talk), a six-person debate about the company vacation, absolute time 274.5 s-285.3 s. Six participants each wore a close-talk microphone; every mic also picks up everyone else in the room. Speakers are identified in the dataset only by aliases — not their real names; NOTSOFAR ships no real names. Do not caption these voices with real names. mic1.m4a .. mic6.m4a each close-talk mic as recorded — its wearer plus the whole room bleeding in. The mic numbering is the job's input order: mic1 CT_20 ("Jerry"), mic2 CT_21 ("Ernie"), mic3 CT_22 ("Carly"), mic4 CT_23 ("Olivia"), mic5 CT_24 ("Sofia"), mic6 CT_25 ("David"). micN-separated.m4a the API's output for that same mic: its wearer, isolated. Same seconds, same job. The job's seventh output — the ambience bed — is deliberately not shipped: over this 10.8 s window it measures -66 LUFS with a 0.003 peak, i.e. silence with faint residue, and a track that needs ~45 dB of gain to be heard does not earn a place in the demo. WHAT WAS DONE TO IT (CC BY 4.0 requires changes to be indicated): a 40-second window (262 s-302 s) of all six close-talk tracks was submitted, unmodified, through the public DUETA API exactly as a customer submits it — six uploads via POST /v1/uploads + PUT .../content, one POST /v1/jobs/dominant-separation naming them in order, polled to success, stems downloaded from /v1/jobs/{id}/stems/{name} (job 8b3159f4-8ce5-46ae-9728-9e0b77d81113, QA account, 2026-09-01, billed $0.33). The pipeline behind that endpoint is sync → prelevel → separate → gate → enhance → level, 24 kHz mono out. The files here are a 10.8 s cut of that job's inputs and outputs, EACH loudness-matched to -20.8 LUFS with a static per-file gain (the mics were recorded at very different gains — mic3's raw window sits near -51 LUFS — and a demo row nobody can hear demonstrates nothing; the relative levels within any one file are untouched), given 5 ms/50 ms edge fades, and encoded to AAC. Nothing else — no manual editing of the audio content, no cherry-modification of the outputs, no path into the pipeline a customer does not have. MEASURED, NOT ESTIMATED. Scores quoted for this clip are computed from the dataset's own word-level ground-truth timings (gt_transcription.json), on this exact 10.8 s window: energy in regions where only OTHER people speak, after vs before separation, normalised on each track's own-speech level. Change in bleed level per mic (negative = bleed removed): mic1 ("Jerry") -1.8 dB mic4 ("Olivia") -31.1 dB mic2 ("Ernie") +7.3 dB mic5 ("Sofia") -0.3 dB mic3 ("Carly") -6.3 dB mic6 ("David") -29.3 dB ALL SIX SHIP, including the rows this window is hardest on (mic2 measures worse after; mic5's own-speech envelope tracks weakly) — the demo is the API's actual output for a real room, weak rows included, not a curated best pair. The clip's overlap is real: two or more people are speaking for 93% of it, three or more for 63%. Those numbers were measured twice, on two independent runs of the same window: once on a direct run of the pipeline driver during candidate scoring, and once on the API job's own downloaded stems. Both measurements are identical to the decimal, and after the same trim/loudness/encode the two runs' files are byte-identical (md5-equal) — the engine is deterministic, so "produced via the API" and "produced by the pipeline" name the same bytes. What ships is the API job's copy. LICENSE. The NOTSOFAR-1 data is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ Attribution is REQUIRED and is carried in the site footer on every page ("recordings from Microsoft's NOTSOFAR-1 dataset, CC BY 4.0") and, in full, here. CC BY 4.0 permits commercial use and modification; it does NOT license publicity/privacy/personality rights (§2(b)(1)) — the dataset's participants were recorded by Microsoft for public release under this license and are identified only by aliases, and these clips add no identification beyond what the dataset itself publishes. If these files are regenerated, re-cut, or replaced with other NOTSOFAR material, keep this notice accurate: the file map, the time range, the processing description and the measured numbers above all describe THESE files, not the directory.