Year End Mega Sale:
30 Days Money Back Guarantee
Discount UP To:
80%
NLP SOLUTIONS

Sorting text at volume,
usually without training anything

Tickets, contracts, claims and reviews arrive as text and have to be classified, read and sent somewhere. A general model with a carefully written prompt handles a great deal of that. When it does, we say so and stop, because that is the smaller invoice.

Get a quote → See the whole service

In short: Classification, extraction and routing across text at volume. This is the area where a general model with a good prompt most often wins outright, so it is the area where we most often argue against training a model of your own. Training earns its cost in three specific situations, and we check for them before quoting one.

The four jobs behind most requests

People describe this work in a dozen ways, but underneath it usually reduces to a small set of tasks. Naming which one you have is the first useful step, because each is measured differently and each fails differently.

Classification
Putting each message, ticket or review into a category, so it reaches the right team or feeds a count somebody reports on.
Extraction
Pulling named fields out of unstructured documents: dates, parties, amounts, policy numbers, renewal clauses, whatever the form needs.
Routing and triage
Deciding where something goes and how urgent it is, with a confidence score attached so uncertain items land with a person instead.
Summarizing at volume
Condensing long threads or dense documents into something a person can read in seconds and decide on without opening the original.

Why we argue against training

Text is where general-purpose models are strongest. They have read enormous amounts of ordinary language, which means the category names in your ticket system and the clause types in your contracts are already familiar territory. Give one a clear instruction, a few worked examples and a strict output format, and it will often match what a trained model would achieve after weeks of work.

Training brings a tail of costs that rarely appears in the pitch. You need labeled data, and someone senior enough to produce labels worth learning from. You need somewhere to host the result, a version scheme, a monitoring setup and a plan for the day performance slips. Then the categories change, because categories always change, and you do it again. If a prompt gets you there, none of that exists. We would rather write the smaller proposal and be useful to you for longer.

Three cases where training does earn its cost

The first is domain language that does not appear in public text. Internal part numbering, coded shorthand your adjusters have used for twenty years, clause conventions particular to your sector, abbreviations that mean one thing in your building and something else everywhere else. A general model guesses at these. A trained one learns them.

The second is volume, where per-call pricing dominates the arithmetic. If you process millions of short documents a month, a small model you host can cost a fraction per item compared with an API call, and the savings pay for the build inside a predictable window. We work that out with your real volumes rather than assuming it.

The third is explainability. Where a decision has to be defended to a regulator, applied identically to two similar cases, and reproduced the same way next year, a fixed model with a recorded version and a documented training set gives you ground to stand on. If that is your situation, it usually arrives with security and compliance requirements that shape the design from the start.

How we decide, and how we measure

Before choosing an approach we build an evaluation set from your real traffic. A few hundred genuine examples, including the awkward, rare and badly written ones, each with an agreed correct answer. That set is the referee. Any approach we try gets scored on it, and you can see the difference between options as a number rather than an opinion.

Building it also surfaces something useful. Have two of your own reviewers label the same hundred items and compare. Where they disagree, no system can do better, because there is no single right answer to learn. That figure is the honest ceiling on the project, and knowing it early stops everyone chasing a target that never existed. Once live, the same set becomes the regression test, and the outputs feed your analytics and reporting like any other structured field.

What we won’t do

We won’t sell a trained model for a prompt-sized problem
If a general model with a written prompt and a proper test set reaches your target, that is the recommendation. It is a much smaller engagement, and we would rather have the reputation than the fee.
We won’t grade a text system on examples we picked
The evaluation set comes from your real traffic, awkward cases included, and your team agrees the correct answers before anything is scored. Otherwise the demo works and the rollout does not.
We won’t route anything sensitive without a human path
Claims, complaints and anything with legal or safety weight keep a confidence threshold and a queue where a person decides. Full automation on those categories is not something we will build.

Do you need a trained model, or a good prompt?

Give us a few hundred real documents and the decision you want made about each one. We will tell you which of the two this is, and we are perfectly happy for the answer to be the cheaper one.

Get a quote →

Frequently Asked Questions

When is training your own model worth it?

In three situations. When your text uses language that does not appear in public data, such as internal codes and sector shorthand. When volume is high enough that per-call pricing outweighs the cost of training and hosting. Or when the decision must be explainable and stable across years. Outside those, a prompt with a real test set usually wins on both accuracy and cost.

How accurate can field extraction get?

It depends on how consistently the documents are written and how consistently your own reviewers agree. If two experienced people disagree on a fifth of the documents, no system will exceed that ceiling, because there is no single correct answer to learn. We measure human agreement first and quote against it.

Can this handle languages other than English?

General models handle major languages well and thinner ones unevenly. We test on your actual documents in each language rather than assuming, since quality can drop sharply on regional forms, transliteration and text that mixes two languages in one sentence. Where a language falls short, we say so before it becomes a rollout problem.

What about confidential documents?

Where the text gets processed is a design decision, not an afterthought. The options include self-hosted models inside your own boundary, contractual limits on data use and retention, and redaction before anything leaves your systems. We pick the arrangement with your legal and security people in the room.