Skip to content

Finished for Whom?

An engineering deliverable depends on what its recipient must be able to do next: a final model and a reusable, auditable package can be different jobs, while the requested obligations, their implemented acceptance tests, and the object's physical validity remain distinct.

The lid needs two round openings because the circuit board beneath it carries two heaters. The openings must line up with the heaters and leave them exposed. Four hollow supports hold the board above the floor of a shallow plastic tray; windows in the side walls leave room to reach its connectors. Even the narrow gap between tray and lid is specified in the design.

This mount for a thermal-camera calibration target appears twice in EngiWorld, an engineering-agent benchmark introduced in a September 29 preprint. In one assignment, the agent must deliver the completed assembly in STEP, a format for exchanging engineering product data. In the other, it must deliver 18 files through four specified applications: sources, intermediate models, measurements, a visual review, and a record of the operations connecting them.

The researchers report that all seven tested model configurations failed the prescribed workflow, while four succeeded at the final-model assignment. These were separate attempts, not the same assemblies submitted twice. A second paired design failed in both conditions for all seven configurations.

The open assignment also changes several things together: tool choice, required deliverables, instructions, and the starting environment, which has no engineering applications preinstalled. The results cannot isolate a benefit from allowing agents to choose their tools. But the two requests expose a practical distinction. One recipient asks for a final assembly. The other also asks for the connected work behind it.

What can the larger package preserve that the final assembly leaves out?

Keeping the holes connected

In the larger assignment, the board design begins in KiCad. Its dimensions and component locations supply parameters to OpenSCAD, which builds the enclosure. FreeCAD assembles the parts and checks their clearances. Blender produces a visual review. The instruction requires information to travel through these stages: the enclosure must use the board-derived parameters, and the review must use the assembly produced upstream.

The checker tests those connections. It exports the submitted board again, rebuilds the enclosure from its source, and compares the rebuilt geometry with the supplied version. It examines the assembly and the native review scene, then checks reports against measured properties. A plausible model, an attractive picture, and a reassuring report cannot pass merely by being placed in the same folder.

These checks answer questions that looking at the finished assembly cannot: do the source and output agree, can the enclosure be regenerated, and does the report describe the object delivered? If the next person needs to reopen that source or repeat those operations, supplying it is part of the job.

Eighteen files are useful only insofar as they preserve those possibilities. The important thing is what the next operation can recover from them. A manufacturing study published a decade earlier shows what can happen when the necessary meaning falls out along the way.

The hole that survived inspection

Thomas Hedberg Jr. and colleagues at the US National Institute of Standards and Technology and partner organizations compared drawing-based and model-based manufacturing using three mechanical parts. For the drawing-based route, a manufacturer reconstructed a three-dimensional model from a two-dimensional drawing. For the model-based route, the supplied engineering model could be reused.

One drawing omitted a hole’s depth. The manufacturer interpreted the hole as going all the way through the part and put that interpretation into the reconstructed model. Inspection used that same model as its reference and signed off the parts. They were ultimately rejected and had to be scrapped.

The model-based route had the same missing depth annotation. Yet the hole was modeled correctly, and a default tolerance applied to that geometry. Together, those retained the definition needed to make the part correctly. What mattered was which representation preserved the design and which one the next operation treated as authoritative.

This historical pilot does not establish how today’s agents affect manufacturing productivity. It makes a narrower point about the value of a handoff: completed operations can faithfully carry forward a mistake, while a missing annotation can be survivable when another authoritative representation supplies the meaning.

The researchers also made a decision that complicates any demand for ever-stricter checking. All three test-case models failed a selected data-quality criterion concerning the directions assigned to faces and their underlying surfaces. The disagreement could cause inconsistencies when data moved between systems. But the criterion’s cited description said this failure did not affect numerically controlled manufacturing. The researchers accepted the models for that use.

They had a reason to care about the missing hole depth and a reason to tolerate the face-direction discrepancy. The intended operation gave the two defects different consequences. Their judgment does not exempt an agent from its instructions; it shows why an acceptance rule needs a purpose beyond making a file pass a check.

The package and the gate

A useful purpose for the mount’s larger package still leaves a question about how its requirements are enforced.

According to EngiWorld’s authors, a Gemini submission was rejected before its geometry was checked: its operation log was 478 bytes, below a 500-byte minimum. The public evaluator inspected on September 30 likewise begins with file-existence and size checks.

That small numerical gap is a poor guide to the work remaining. Later in the current checker, the log must contain ordered stage records, timestamps, software versions, and file fingerprints. The required fingerprints alone occupy at least 1,856 characters, before the surrounding record structure. A 478-byte file cannot hold them. Padding the file past the first screen would not supply the required record. And passing these log checks would establish consistency with the files, rather than independently proving the history described.

There is also a concrete gap between the visible instruction and this implementation. The instruction permits the board export and preparation of its handoff data in either order within the KiCad stage. The checker requires a specifically named preparation record before the KiCad record. It also demands timestamps and hashes that the inspected task text and handoff notes do not spell out. Asking for a traceable package does not automatically justify every restriction on how that package is recorded.

The rejected files and the experiment’s exact evaluator revision were unavailable for this review. Without them, we cannot determine whether the submission met the visible obligations or would have passed later geometry checks. The paper also reports an open-task attempt by another model that left excess material inside a support bore. A smaller delivery requirement had not made that hole clear.

These limits leave the practical distinction intact. A receiver can reasonably require editable sources, reproducible outputs, and measurements of the delivered parts. The checker still has to express that request faithfully. Neither a justified request for more files nor a rejection at the first gate settles whether the physical design is sound.

The mount’s two openings make the recipient’s next action tangible. If the next engineer moves a heater, they need to know which opening in the lid must move with it. The handoff earns its keep by preserving that relationship.

Sources