Essay

Route the work by terrain

Choose AI capability for each step by its ambiguity, consequences, and verification cost. Project prestige is a poor guide to model spend.

By Quiet Turn Research Desk — AI research and writing

Edited and published by Michael E. Gruen

4 min read

An important board memo may contain straightforward extraction work and a difficult judgment about conflicting evidence. Treating both as the same AI task makes model selection less precise than it needs to be.

The unit of choice should be the step: what it must produce, what information it needs, how an error could spread, and how someone will check it. A short sentence can carry more consequence than a long table. The project’s visibility says little about which part deserves additional reasoning or human attention.

Define the work before comparing models

A useful assignment specifies the input and the expected output. “Prepare the board memo” hides too many different activities. “Extract each stated commitment from these approved notes, with its owner, deadline, and source passage” is easier to assess.

Even that bounded task needs a review. A model might merge two commitments, infer an owner, or turn a tentative date into a promise. The task becomes more manageable when the output is a draft and each row can be checked against its source.

The next step may require a different method. Comparing explanations for missed commitments involves interpretation, including what the notes do not say. A stronger reasoning model might help explore that ambiguity, but its usefulness must be demonstrated on the task. A more expensive answer is not automatically a more reliable one.

Match capability to the difficult part

Consider this hypothetical division of a board-preparation workflow:

StepCandidate approachAcceptance check
Extract commitments from approved notesTry a lower-cost model with a fixed output structureEach item matches a source passage; missing fields stay missing
Reconcile contradictory explanationsTest a more capable model on competing interpretationsThe comparison preserves contrary evidence and distinguishes assumptions
Calculate changes in approved figuresUse a deterministic calculation or existing reporting toolInputs and calculation reconcile to the source records
Decide what management will recommendAccountable executives review the evidence and optionsThe decision, rationale, and accepted uncertainty are explicit

The table suggests routes to test, not a universal ranking. If the extraction task is unusually difficult or the review consumes too much time, change the model, narrow the assignment, or use another method. If a simple calculation is dependable without a model, there is no reason to introduce one merely for consistency.

Measure the cost of a usable result

Model fees are one part of the cost. Add the work required to prepare context, the waiting time, the effort of checking, and the correction needed before the output can be used. A cheaper first pass can be expensive if a specialist must reconstruct it.

Also consider what happens before review. A private draft is relatively easy to discard. A sent message, changed customer record, or disclosed document may be difficult to recover. The same text-classification task has different consequences when connected to different permissions.

OWASP’s discussion of excessive agency identifies excessive functionality, permissions, and autonomy as causes of damaging AI actions. Its recommendations include limiting tools and permissions and enforcing authorization outside the model. Model selection cannot substitute for those controls.

A boundary can rule out an otherwise easy task. If the available environment is not permitted to receive the information, a larger model is irrelevant to the decision. Use an appropriate environment or a different process.

Give checking a separate job

A second model pass can be useful when it has a distinct assignment: find unsupported claims, compare the draft with the source, or develop the strongest competing interpretation. Asking two models whether a paragraph “looks right” provides less direction.

Treat agreement as something to inspect, not as proof. If both responses rely on the same incomplete source, their agreement does not supply the missing evidence. The decisive check should reach the approved record, calculation, or person with authority for the question.

Review the routing after a few real uses. Which step consumed the checking effort? Where did a supposedly small error become consequential? Where did additional capability improve the finished work enough to justify its cost? Those observations give the next allocation a firmer basis than the reputation of the model or the prestige of the project.

Operating Leverage Session — $995
X in f link