Access
Two ways in.
A free sample you can browse before talking to anyone, and a commercial license negotiated per buyer. Questions buyers actually ask are answered below.
Questions buyers actually ask
Asked before the first call.
Can we train a commercial model on this data and ship the product?
Yes. The consent form patients sign lists the permitted uses, and “commercial development or licensing involving dental AI, CAD/CAM, robotics, workflow analytics” is one of them, alongside “AI model development” – quoted in full in the rights section. Scope, term and exclusivity are set in a Data Use Agreement negotiated per buyer; the DUA prohibits re-identification and unauthorized redistribution.
What exactly arrives, and in what formats?
One tree per case: the H.264 dual-view master with a burnt-in timer, narration as separate FLAC and M4A, every annotated keyframe as a full-resolution JPEG, the six-layer annotation JSON with per-label provenance, the transcript as JSON and SRT, the blind vision-only baseline with its comparison report, the compliance summary in Markdown, HTML and PDF, and a SHA-256 checksum for every file. The released sample already carries a LeRobot v2.1 export; COCO-style and Parquet conversions on request.
How is this different from Cholec80 or YouTube-sourced corpora?
Three ways. Specialty: a systematic review counts 16 public dental datasets – 10,450 still images, zero intraoperative video; surgical-video corpora live in laparoscopy and open surgery, not dentistry. The audio: the operating surgeon narrates as he works, and every quote is a timestamped L0 label – scraped video carries no controlled narration at all. The rights: research corpora are released for research; YouTube-sourced material has no consent chain to grant commercial rights from. The comparison table puts the three side by side.
Is there a sample we can evaluate before talking to you?
Yes – a CC BY 4.0 sample of the annotated corpus is live on Hugging Face: 1,009 annotated keyframes, the six-layer JSON, the surgeon's transcript, a LeRobot v2.1 export and a 7-minute annotated demo, browsable in the dataset viewer before you download anything. The link is in the access section below.
What about patient consent and de-identification?
Cases in the line are recorded under prospective written consent – signed before the camera is switched on, on the WCG-approved form quoted above. De-identification follows HIPAA Safe Harbor, 45 CFR §164.514(b)(2), is documented per case after a frame-by-frame visual audit, and PI sign-off precedes every external release. Identifying passages are cut, not blurred.
Do the frames carry bounding boxes or segmentation masks?
No, and the page does not pretend otherwise. L2 records which instruments are in the frame – a per-keyframe list, each entry carrying its confidence and its source (seen in the image, stated by the surgeon, or both) – not where they sit in the pixels. There are no boxes, masks or coordinates in the released labels. What lets you trust the lists is shipped alongside them: the blind vision-only pass over the same frames, with the comparison report.
We need a procedure class you have not released. Can you record it?
Yes – that is what a production line is for. The protocol covers the whole specialty: implantology, dentoalveolar surgery, orthognathic and reconstructive work, and related facial procedures. Recording, narration and the six-layer annotation pipeline apply to any class on that list – commission a procedure-class dataset.
Access
Two ways in.
Open sample – free
A CC BY 4.0 sample of the annotated corpus, live on Hugging Face: 1,009 annotated keyframes, the six-layer annotation JSON, the surgeon's transcript, a LeRobot v2.1 export and a 7-minute annotated demo. Browse it in the dataset viewer before downloading anything.
Commercial license – the corpus and the line
The flagship annotated case with all six layers and full documentation today; two further annotated cases – a maxillary full-arch and a maxillary sinus surgery – as they clear release; new cases entering the pipeline continuously as the line records year-round; commissioned procedure-class datasets to specification. DUA, individual terms – conversations happen over email and calls, not forms.
What we are looking for
Robotics and surgical-AI teams who need language-grounded clinical data · commissioned datasets for specific procedure classes – from wisdom teeth to orthognathic · research collaborations on weak/scalable labeling – a second pass labeled without the surgeon's audio gives a ready-made benchmark · dataset partnerships with implant and navigation manufacturers · grant and consortium partners for scaling annotation to the full recording volume.
Request access
Tell us who you are and how to reach you. Conversations about scope, term and exclusivity happen over email and calls, not forms – this is just the way in.