↑ Back to the surface
40 mLegal, e-discoveryShipping

ReefML

A classification API that reads a legal matter and scores every document for relevance, so review teams read the useful few percent instead of all of it.

2M
documents per matter
87%
recall, 22K doc benchmark
300K
docs in production validation

Document review is the most expensive part of litigation, and most of what gets read turns out to be irrelevant. ReefML learns from a small labelled set and scores the whole corpus responsive or non responsive, so reviewers work a ranked queue instead of a flat list. I own strategy and delivery, and I wrote the TF-IDF vectorizers that turn raw matter documents into the features the classifier actually scores.

Detail

  • Built the TF-IDF vectorizers feeding the scoring and iterative training paths.
  • Up to 2 million documents can be processed for a single matter.
  • Stakeholder alignment for the production testing path is complete.
  • Benchmarked on quality against competitor tooling, not just on latency and cost.

Ranked review queue

DOC-04812
0.97
DOC-01199
0.94
DOC-07340
0.88
DOC-02615
0.71
DOC-09004
0.34
DOC-05527
0.12
ResponsiveNon responsive

My role

Product leadData scientist / developerQA

Built with

PythonSpark MLTF-IDFDatabricksFastAPIDockerAzure