Rohan Paul
@rohanpaul_ai
A frontier model reading a contract will sometimes cite a clause that isn't there.
Overmind fine-tunes a small open model on your own production traces, then scores it against your current model on evals built from those same real tasks.
Their published comparison is Qwen3.5 9B tuned through Overmind against GPT5.6 Luna. The company reports 20 to 30x fewer phantom clauses in legal contracts and 7x better accuracy at quoting a clause word for word.
Here the observability layer and the training layer share the same data. The traces that show you where the agent fails become the dataset, and the evals are built from real tasks rather than a public benchmark.
The output is a smaller open model with weights you own. You can host it on Overmind or run it yourself.