Evaluation for all three scenarios, starting with contract review: expert-labeled ground-truth sets, scoring rubrics, side-by-side model and prompt comparisons, and a failure taxonomy.
A legal AI agent built for the back office
No dataset and no idea to start with. Interviews with practicing lawyers showed who would trust AI and who would not, and Lawrify was built for the people who already check every document.
- Mar 2021 — Dec 2022
- 15-person team
- Legal domain
$150 per seat
proof of concept
from ~60s
Starting point
Lawrify started from zero: no dataset, and no clear idea of what to build. Brainstorming did not produce anything worth building.
So the work started in the market, with interviews with practicing lawyers about what their work actually is.
My role
Lead ML/AI Product Manager at Yandex, leading a 15-person team building Lawrify 0→1.
Translated law-firm workflows into product requirements and built the evaluation stack end to end.
What I did
- Found who would trust AI
Law firms split into a front office, which presents the case and owns the outcome, and a back office, which prepares cases and searches data and precedents. The front office would not risk its reputation on an AI mistake. The back office already double-checks everything, so reviewing AI output fits its habits. Lawrify was built for the back office.
- Picked three scenarios
Find similar legal cases inside and outside the organization. Find and extract data such as citations of specific laws. Review contracts for vague clauses, risks and typos.
- Found the paying audience
Back offices of large law firms and small boutiques became the core users, on a price of $150 per seat.
How we measured it
Recall over precision: the agent should not miss a risky clause.
The back office reviews every flag anyway, so the agent was tuned to find more rather than miss.
Evaluated both the datasets the model produced and the outcome of the agent’s full task.
Lawyers double-checked false positives and false negatives in flagged clauses.
Human evals drove the quality and latency trade-off, cutting response time from about 60 seconds to 15–20.
What changed
A working proof of concept in four months, the first back-office lawyers using it, and around $10K in revenue at $150 per seat. Response time dropped from about 60 seconds to 15–20.
Summary
Reputation decided the wedge. The lawyers who own outcomes would not bet them on AI; the ones who already check every document would, and they review its output by habit.