A smaller pretrained model with constrained inputs and explicit review can outperform a larger dependency once production cost, failure containment, and data boundaries are included.
Model choice tends to get argued on benchmark leaderboards and decided on production invoices. Neither is the right place on its own.
A smaller model you can host inside your own boundary removes a category of problem rather than a percentage point. No per-token cost that scales with your success, no vendor rate limit during your busiest hour, no data-residency conversation with legal, no capability that quietly changes on a Tuesday.
It also gives things up, and pretending otherwise is how the decision goes wrong. You lose headroom on the hardest inputs and you take on operational work you did not have before. Someone now owns patching, scaling, and the GPU bill.
The way through is to measure your own tasks instead of the benchmark: the real prompts, the real documents, the real failure cases. Often the smaller model only loses on inputs that were already going to a person for review, which means it was never the deciding factor.