What a BMT is: a complete guide to validating solutions on your own documents

A BMT (benchmark test) is the validation procedure in which an organization considering adoption runs several candidate solutions against its own real business documents under identical conditions and predefined criteria, compares the results, and decides on that basis.

Why a BMT is necessary

Published scores and scores on your documents are different numbers.

Accuracy is the figure exchanged most often in evaluation meetings. Yet the same product returns 99 percent on one document set and settles in the low eighties on another. In actual validations, table structure restoration was confirmed above 99 percent on filing forms while the same measure fell below 80 percent on administrative documents where multi-level headers and merges overlap. Document classification shows the same spread: 97 percent on one financial document set, 89.1 percent on a public-sector set.

That gap comes from the character of the documents, not the ranking of the products. Deciding on published benchmark scores or the accuracy quoted in a proposal therefore means meeting a different reality once you are in production. A BMT moves that risk forward, before adoption.

What a BMT actually is: how it differs from a PoC

BMT and PoC across three axes

First, the subject differs. A BMT covers several candidate products, throwing the same documents at each under the same conditions so the results line up side by side. A PoC covers a single solution and checks whether it works once attached to the real business flow.

Second, the objective differs. A BMT is about selection: which product does better on our documents. A PoC is about design: how the business rules, exception handling, and system integration should be configured.

Third, the metrics differ. A BMT uses measures that support product comparison, such as document classification accuracy, field-level exact match, and table structure similarity. A PoC uses operational measures such as handling time per case, the share of fields an operator touched, and the proportion of cases routed as exceptions. In one financial project, the BMT set an exact match target of 90 percent and measured classification and field extraction, with operating thresholds fixed at a later stage.

The two are not alternatives. Narrowing candidates with a BMT and settling the operating design with a PoC is the usual order, though a simple document set with one obvious candidate is sometimes handled in a single combined exercise.

Six design principles that make BMT results trustworthy

Six principles that hold in a real exercise

First, sample the evaluation documents to match the operational distribution. Gathering only the clean documents means the score drops after go-live; gathering only the hard ones means no product distinguishes itself. Mix document types and conditions in proportion to actual intake. One financial project organised an initial set of 46 document types and selected 50 for evaluation; one public-sector project drew 302 assessed elements from eight documents.

Second, build the ground truth first. Without a human-confirmed answer for each field there is no way to compute a score. This is the most time-consuming part of a BMT, and cutting corners here destabilises every comparison that follows.

Third, define the metrics in advance. Accuracy means a different calculation at every vendor. Settle in writing whether you are measuring exact match or fuzzy similarity, whether tables are judged on structure similarity, and at what unit classification is counted, before anything runs.

Fourth, include exception documents without fail. Stamps over tables, fax headers stacked several deep, blank fields that require a judgment call — these are what separate products. On clean documents alone, most of them score alike.

Fifth, hold conditions identical. The same files, the same resolution, the same field definitions. Tuning the extraction criteria more carefully for one candidate turns the result into a measure of preparation time rather than product capability.

Sixth, evaluate the deployment environment alongside performance. Equal accuracy is irrelevant if internal installation is impossible or the accelerator footprint is unreasonable. Put the performance figures and the deployment requirements in the same table.

How a BMT is run

Finish document collection and field definitions first

The first deliverable of a BMT is not a result but a document list and a field definition sheet. Which forms, how many of each, which fields from each form and in what format — settling those creates the basis for comparison.

It is worth organising the classification scheme at the same stage. One financial project built a classification code scheme of 747 entries while organising its document set, then moved on to extraction validation. Measuring extraction performance without defined classification means missing the bottleneck that will appear in production.

Separate measurement from interpretation

Measure mechanically; interpret separately. A low score on a field may reflect a limitation of the product, an ambiguous definition in the ground truth, or the absence of any basis in the source document itself.

In practice, a blank cell may be not applicable, an omission, or a scanning loss, with nothing in the document to decide between them. Counting such fields as errors turns a product evaluation into a document quality evaluation. Classify undecidable fields separately and set the handling policy as a business rule.

Running a BMT in the Korean environment

The first obstacle in Korean financial and public-sector BMTs is where the documents can be processed. Network separation prevents evaluation documents from leaving the organization, so confirm first whether candidate products can be installed temporarily inside the network to run the exercise. A product that cannot meet this condition cannot be evaluated at all, regardless of performance.

Check the format conditions too. Confirm that Korean word processor files, government forms, and fax-received documents are in the evaluation set. Figures measured on published sample documents and figures measured on your own intake queue come out differently. That difference is precisely why the BMT exists.

Frequently asked questions

They are useful as reference. Public evaluation sets are composed differently from your documents, so the ranking does not reproduce directly. Base the adoption decision on results measured against your own documents.

It depends on the diversity of the set. One financial project ran with 50 form types; one public-sector project drew 302 assessed elements from eight documents. Reflecting the actual intake distribution matters more than the count.

Yes. Without confirmed answers there is no score to compute and no comparison to make. It is the most laborious part of a BMT and cannot be skipped.

It depends on the work. Lead with field-level exact match where values are entered into a system, table structure similarity where tables are the substance of the document, and document classification accuracy where intake is what you are automating.

Yes. Installing candidate products temporarily inside the network allows validation without documents leaving the organization.

Usually there is one more stage. Business rules, exception handling, and system integration are settled in a validation phase before the transition to production.

Related terms