Method
Every quality claim has a number behind it.
Annotation quality is measured against a blind vision-only pass over the same frames, and the comparison ships with the data so you can recompute it.
Measured, not asserted
Every quality claim on this page has a number behind it.
Inter-annotator agreement is measured on 117 double-annotated overlap pairs. And a full second annotation pass was performed blind – without the surgeon's audio – to quantify exactly what expert narration contributes. That control arm is not a one-off: it runs on every case in the line.
| Signal | Vision-only pass | With the surgeon |
|---|---|---|
| Surgical phase identification | agrees on 54.4% of keyframes | phase agreement 1.00 between annotators |
| Instrument presences found | 3,685 – misses ~22% | 4,707, each with confidence & source |
| Clinical context (L5) depth | mean 112 chars | mean 202 chars – 80% richer |
| Surgeon intent (L0) | does not exist – no audio | 1,095 timestamped statements |
Definitions: 1.00 = phase-label match on all 117 double-annotated overlap pairs; 54.4% = the blind pass vs the audio-informed reference across all 1,056 keyframes; −22% = 3,685 instrument labels found blind vs 4,707 with the audio.
Mean instrument-list Jaccard across 117 overlap pairs on a 39-term multi-label task – and every residual disagreement stays inspectable in the released data.
Structured question rounds with the operating surgeon – three answered and already applied to the labels, a fourth in flight – resolved terminology, the drill protocol and component naming.
Of surgeon quotes traceable to source audio. ASR errors on clinical terms are corrected against surgeon review, never accepted silently.
Cases through the blind control arm. On the maxillary full-arch, reasoning came out 137% longer and clinical context 505% longer once annotators had the audio. On the sinus-surgery case, 45% of frames could only be named from the surgeon's speech – the sinus interior is never visible to a camera – so every phase there carries a label_source field: speech, image or both.
Modalities & scale
Built as a production line – capacity for hundreds of cases a year.
The contributing practice performs 40–50 full-arch implant cases per month across four IRB-approved locations, under one active protocol and one consent workflow. Recordable volume is not the constraint here; annotation throughput is, and we are scaling it toward ~200 annotated procedures per year.
Shipping now available
The line never stops: the released flagship, a 1:27:22 maxillary full-arch (1,735 keyframes, 100% annotated) and a maxillary sinus surgery (648 keyframes) clearing release, and further cases in recording and annotation behind them. Each ships as a de-identified dual-view master – dual native-4K cameras, synchronized transcript, six annotation layers, full de-identification audit trail. 4K source access is a license-tier conversation.
In the pipeline annotated, clearing release
A maxillary full-arch case – 1:27:22, 1,735 annotated keyframes – and a maxillary sinus surgery – 648 keyframes – are fully annotated and moving through release clearance behind the shipped case. New recordings enter the line continuously.
Scales beyond one surgeon protocol design
The IRB package is written for participating clinicians, not a single operator: four approved locations today, and new surgeons join under the same protocol after required training, documentation, site activation and any WCG-required personnel update or approval. Capacity grows by onboarding clinicians – the consent workflow stays identical at every site.
Recording is routine in-house OR
Surgery and anesthesia run in the practice's own operating rooms – general-anesthesia permit #GA 1446 for in-office IV sedation, plus an elective facial cosmetic surgery permit for the adjacent procedure classes. The rig is set up where the operating happens, so any workday can be a recording day.
Rig roadmap validated in-house
RGB-D depth stream – depth already recorded on the pilot case, production camera selection in progress. Navigation telemetry planned as the kinematics-analog layer.
Delivery formats
Annotation JSON with per-label provenance; a LeRobot v2.1 export already ships in the released public sample; COCO-style and Parquet conversions on request. Documentation to the level of a reproducible pipeline, including its failure modes.
Submitted to Open-H-Embodiment v2 – the healthcare-robotics data initiative led by Johns Hopkins, TU Munich and NVIDIA (Track B – annotations), August 2026.
Procedure coverage
One protocol covers the whole specialty.
The IRB protocol is titled for its true scope – adult dental, oral/maxillofacial, and related facial reconstructive/cosmetic procedures. Full-arch implantology is the flagship procedure class, not the boundary: every class below is performed at the practice and recordable under the same active protocol and the same per-subject consent workflow. Two of the classes below already hold annotated cases.
Implantology flagship
All-on-4 / All-on-X full-arch reconstruction · single-tooth implants · zygomatic implants · bone grafting · computer-guided surgery. Annotated cases on both jaws are in the line.
Dentoalveolar surgery recorded
Wisdom-teeth and complex extractions · oral pathology · maxillary sinus surgery – the first annotated case of this class is in the line · pre-prosthetic surgery
Orthognathic & reconstructive
Le Fort osteotomy · BSSO · genioplasty · facial trauma reconstruction · TMJ surgery
Related facial procedures
Facial cosmetic surgery within the protocol's approved scope – rhinoplasty, face lift – performed under the practice's elective facial cosmetic surgery permit
The end product of the flagship class: a screw-retained full-arch prosthesis on multi-unit abutments. The surgery that places it is what this dataset annotates, frame by frame.
Need a dataset for a specific procedure class? Recording, narration and the six-layer annotation pipeline apply to any of the above – commission a procedure-class dataset →