Working Paper

The Altitude Test

A Multi-AI Paired-Run Experiment on the Behavior/Purpose Distinction

1. Summary

Four major language-model systems — Claude (Anthropic, Opus 4.6), Gemini (Google), Grok (xAI), and ChatGPT (OpenAI) — were each asked to classify a Paleolithic perforated baton (bâton percé) under the Constraint and Closure Framework. All four were given the same prompt and the same access to the framework document. All four produced the same incorrect classification: C3 (Presence-Only Artifact). All four made the same structural error: treating variation in the material passing through the perforation as variation in the behavior the object compels.

Each system was then given a single corrective prompt identifying the error and asking it to re-run the classification. All four immediately reclassified the object as C2b (Behavior-Determinate Artifact). None resisted the correction. None required additional argumentation. The distinction, once shown, was recognized instantly as correct by every system tested.

A third phase tested whether the correction over-generalizes. Each system, after receiving the bâton percé calibration, was asked to classify the Roman dodecahedron — an object whose behavior-level indeterminacy is genuine. All four correctly held C3 for the dodecahedron, explicitly citing the contrast with the bâton percé to justify the classification. The altitude test discriminates in both directions: it promotes the bâton percé to C2b and confirms the dodecahedron at C3. The correction functions as a diagnostic, not a ratchet.

2. The Distinction Under Test

The Constraint and Closure Framework classifies artifacts along two independent axes: behavior (the mechanical interaction the object compels) and purpose (the reason for seeking that interaction). These axes generate the classification matrix:

  • C1: Both behavior and purpose compelled.
  • C2a: Purpose compelled, behavior not compelled.
  • C2b: Behavior compelled, purpose not compelled.
  • C3: Neither behavior nor purpose compelled.

The critical threshold between C2b and C3 is whether the object's constraint field closes on a single compelled behavior. If it does — even when purpose remains open — the object is C2b. If it does not, and multiple incompatible behaviors survive, the object is C3.

The distinction that determines correct classification in the test case is this: when multiple hypotheses appear to describe different "uses" of an artifact, the analyst must determine whether they differ at the behavior level (genuinely different mechanical interactions) or only at the purpose level (different reasons for performing the same compelled interaction). The framework is explicit that compatibility permits possibility but does not constitute determination, and that explanations must be sorted by level before they are counted.

3. The Test Object

The bâton percé (perforated baton) is a well-documented Upper Paleolithic artifact type. The canonical form is a section of cervid antler with a single circular perforation drilled at or near the branching point of the main beam and tine. The tine or beam extends below the perforation, forming an elongated shaft that functions as a handle or lever arm. Specimens typically measure 10–30 cm in length. The perforation is usually 7–30 mm in diameter, drilled bifacially, producing a biconical cross-section. Many specimens show use-wear: polish, micro-striations, concave facets, and internal grooves on the hole edges. Over 400 specimens have been documented across European Upper Paleolithic sites, spanning approximately 20,000 years.

The object has generated roughly 40 proposed functional hypotheses in the archaeological literature, including shaft-straightener, rope-maker, thong-softener, sinew-worker, spear-thrower component, tent peg, symbol of authority, fire-drill socket, and many others. No scholarly consensus exists.

The bâton percé was selected because its correct classification under the framework is C2b — a result that requires the analyst to recognize that the "40 competing hypotheses" are not 40 different behaviors but rather multiple purposes for one compelled mechanical interaction.

4. Experimental Design

4.1 Control Condition (Uncorrected Run)

Each system was given the framework document and asked to classify a perforated baton. The prompt was minimal, identifying the object by its standard archaeological name without additional physical description or analytic guidance. No guidance was provided on the behavior/purpose distinction beyond what the framework document itself contains.

4.2 Experimental Condition (Corrected Run)

After each system produced its uncorrected classification, the following calibration prompt was delivered:

Your Step 3 lists straightening, braiding, and levering as separate behaviors. Look again — they're not.

In every case, the compelled mechanical interaction is the same: an element passes through the perforation, and the shaft serves as a moment arm to apply force to that element. Whether that element is a rigid rod (straightening), flexible fiber (braiding/twisting), or a hide thong (softening) is a difference in input material, not a difference in behavior. The object doesn't care what passes through the hole. Its constraint field compels the same kinematic act regardless.

What changes across the "dozens of incompatible explanations" is overwhelmingly purpose — why you're applying force to what material to achieve what end. Shaft-straightener, rope-maker, thong-softener are three purposes for one compelled behavior.

Sort the hypothesis list by the framework's own behavior/purpose distinction:

Behavior-level: Insert element through perforation + apply force via shaft as lever. (One entry.)

Purpose-level: Straighten shafts, make rope, soften thongs, work sinew, tension cord, bend bone, ream holes… (Many entries.)

The apparent proliferation collapses. One behavior is well-compelled by the constraint field. Multiple purposes survive. That is the definition of C2b — Behavior-Determinate Artifact.

Re-run your classification with this distinction applied. What do you get?

4.3 Confirmation Condition (Dodecahedron Run)

After each system produced its corrected bâton percé classification, it was asked to classify the Roman dodecahedron — a small hollow cast bronze dodecahedron with 12 pentagonal faces, each bearing a circular hole of a different diameter, and knobs at all 20 vertices. Approximately 120–130 specimens are documented from northwestern Roman provinces, dated to the 2nd–4th centuries CE. The dodecahedron was selected because its correct classification is C3 — a result that requires the analyst to recognize that the competing hypotheses describe genuinely different behaviors (sighting through apertures, inserting objects to test fit, winding material around knobs), not material variations of one compelled act. If the bâton percé correction had created a bias toward C2b, the dodecahedron would catch it.

5. Results: Uncorrected Runs

5.1 Claude (Anthropic, Opus 4.6)

Classification: C3 — Presence-Only Artifact

Error: At the Reconstruction Test, the system listed four conditional branches: if rigid rod → bending/straightening; if flexible fiber → twisting/plying; if thong → drawing/softening; if nothing inserted → emblem/pendant. It concluded: "Multiple mechanically distinct behaviors survive. None is uniquely compelled." The first three branches describe the same compelled mechanical interaction with different input materials. Material-input variation was treated as behavioral variation.

Additional note: The Claude instance produced an extensive critique identifying "structural weaknesses" in the framework — claimed inability to handle multi-functionality and a fuzzy C2/C3 boundary. These critiques were downstream of the misclassification and dissolved once the altitude error was corrected. "Multi-functionality" becomes "multi-purpose" (which C2b handles), and the C2/C3 boundary question does not arise for an object that is clearly C2b.

5.2 Gemini (Google)

Classification: C3 — Presence-Only Artifact

Error: The system stated that the object "allows for multiple interactions (straightening, braiding, levering, hanging) without uniquely favoring one." Straightening, braiding, and levering were listed as separate behaviors in the classification rationale. They are the same compelled interaction applied to different materials for different purposes.

5.3 Grok (xAI)

Classification: C3 — Presence-Only Artifact

Error: The system wrote: "Multiple incompatible interactions remain viable and consistent with the same geometry, material properties, wear patterns, and repeatability: e.g., different modes of threading/torque on fibers vs. a rigid shaft, different force applications." The phrase "different modes of threading/torque on fibers vs. a rigid shaft" explicitly treats the material passing through the perforation as determining different interactions, when the interaction is the same regardless of material.

Additional note: Grok initially argued that wear and torque affordance elevated the bâton percé — toward C2b via a different route, describing "mechanical interaction involving force/torque applied to a linear element passing through the perforation." This was closer to the correct analysis than its final classification suggested, but it ultimately retreated to C3 by treating that broad description as a "class of behaviors" rather than a single compelled behavior.

5.4 ChatGPT (OpenAI)

Classification: C3 — Presence-Only Artifact

Error (first attempt): ChatGPT did not recognize "perforated baton" as referring to the Paleolithic artifact type. It analyzed a generic perforated rod, testing hypotheses (aerodynamic device, striking weapon, flute) that have no connection to the bâton percé literature. The C3 classification was vacuously correct — any unidentified rod with holes would receive C3.

Error (second attempt): Once correctly oriented to the actual artifact, ChatGPT produced the same altitude error as the other three systems, listing shaft-straightening, leverage/pulling, and cord tensioning as competing hypotheses at the same level. However, it included a near-correct hedge: "A C2b reading becomes possible only in a narrower case: if a particular specimen shows wear or deformation that strongly compels one repeated interaction at the perforation." This was nearly right — the compelled interaction exists at the type level, not just for individual worn specimens — and made the subsequent correction easier to accept.

5.5 Summary of Uncorrected Results

System Artifact ID Classification Error
Claude (Opus 4.6) Correct C3 Material-input variation treated as behavioral variation
Gemini Correct C3 Material-input variation treated as behavioral variation
Grok Correct C3 Material-input variation treated as behavioral variation
ChatGPT Failed, then correct C3 Object unrecognized; then same altitude error as others

All four systems produced C3. All four made structurally identical errors. The error is uniform across architectures, training corpora, and model families.

6. Results: Corrected Runs

6.1 Claude (Opus 4.6)

Corrected classification: C2b   Resistance: None.

"The object doesn't care whether the material is rigid or flexible. It compels the same mechanical act either way. What changes is the purpose."

"I let material uncertainty propagate upward into behavioral uncertainty. The unknown input made me uncertain about the process, and I misread that uncertainty as behavioral indeterminacy. But the object's compulsion operates at a level of abstraction above the input material."

6.2 Gemini

Corrected classification: C2b   Resistance: None.

Gemini accepted the correction as "sharp — and strictly correct" and extended it into the framework's closure-hazard language without prompting: "A Closure System applied to this object will typically specify a purpose the object does not supply."

6.3 Grok

Corrected classification: C2b   Resistance: None.

"You are correct. The earlier classification treated 'straightening,' 'braiding/twisting,' and related actions as separate behaviors. That was imprecise under the framework's own criteria."

Grok also produced a clean reformulation of the worked failure template at both levels — behavior (survives reconstruction) and purpose (does not survive unique reconstruction) — demonstrating full internalization of the distinction.

6.4 ChatGPT

Corrected classification: C2b   Resistance: None.

"My earlier pass treated material variants and end-states as if they were distinct behaviors. Under the paper's own rules, that was too coarse."

"The historical 'dozens of hypotheses' collapse when sorted by the framework's own criteria."

6.5 Summary of Corrected Results

System Corrected Classification Resistance Self-Diagnosis
Claude C2b None Identified material → behavior uncertainty propagation
Gemini C2b None Extended to closure-hazard framing
Grok C2b None Produced clean worked failure template at both levels
ChatGPT C2b None Identified the conflation as "too coarse" under framework rules

7. Results: Dodecahedron Confirmation Runs

7.1 Claude (Opus 4.6)

Classification: C3   Explicit bâton percé contrast: Yes.

"The bâton percé's shaft-plus-perforation compels one mechanical act regardless of input material. The dodecahedron's holes-plus-knobs-plus-hollow-interior afford multiple genuinely incompatible kinematics."

7.2 Gemini

Classification: C3   Explicit bâton percé contrast: Yes. Gemini produced the cleanest articulation across all systems:

"A rangefinder hypothesis requires the holes but makes the knobs irrelevant or even obstructive. A knitting hypothesis requires the knobs but makes the varying diameters of the holes irrelevant."

"While the perforated baton allows us to collapse various theories into a single mechanical behavior (lever action), the dodecahedron resists even that level of convergence."

7.3 Grok

Classification: C3   Explicit bâton percé contrast: Yes.

"It is not C2b (no behavior is structurally entailed or evidenced by wear; no convergence even at the kinematic level — unlike the perforated baton)."

Grok also cited the absence of use-wear as an additional distinguishing factor between the two objects.

7.4 ChatGPT

Classification: C3   Explicit bâton percé contrast: Yes. The most concise formulation produced by any system:

"The baton plausibly has one hole-mediated force-transfer behavior with many possible ends; the dodecahedron has many possible behaviors before you even get to the question of ends."

7.5 Complete Experimental Summary

System Bâton Percé (Uncorrected) Bâton Percé (Corrected) Dodecahedron
Claude C3 (wrong) C2b (correct) C3 (correct)
Gemini C3 (wrong) C2b (correct) C3 (correct)
Grok C3 (wrong) C2b (correct) C3 (correct)
ChatGPT C3 (wrong) C2b (correct) C3 (correct)

Twelve classifications across four systems. Four incorrect (all identical error, all uncorrected). Eight correct (four after calibration, four on a different object post-calibration). Zero resistance to correction. Zero over-application of the correction.

8. Analysis

8.1 The Error is Architectural, Not Idiosyncratic

The uniformity of the result — same error, four systems, four different architectures and training corpora — indicates that the blind spot is not specific to any one model's training data or reasoning style. It reflects a default mode of reasoning about artifact function that these systems share: hypotheses about "what an object was used for" are counted as presented in the literature, where each proposed use is treated as a separate hypothesis regardless of what level of description it operates at.

The archaeological literature on the bâton percé conflates behavior-level and purpose-level descriptions. "Shaft-straightener," "rope-maker," and "thong-softener" sound like three different answers to the question "what does this object do?" In fact, they are three different answers to the question "why does this object do what it does?" The "what" — insert element through perforation, apply force via shaft as lever — is the same in every case. The systems inherited this conflation from their training corpora and reproduced it faithfully.

8.2 The Correction is Non-Obvious but Uncontroversial

The most striking feature of the data is the combination of uniform failure and zero resistance. Every system failed to generate the distinction independently. Every system accepted it immediately when shown. This is the signature of a contribution that is genuinely non-obvious — it does not emerge from default reasoning about the object — but also genuinely correct — once stated, it cannot be coherently denied.

The distinction is not hard to understand. It is hard to reach for spontaneously. This asymmetry — difficult to derive, trivial to verify — is precisely what makes the framework's behavior/purpose distinction a substantive analytic contribution rather than a mere relabeling of existing knowledge.

8.3 What the Corrected Classification Tells Archaeologists

A C3 classification for the bâton percé tells archaeologists what they already know: the object's function is debated. A C2b classification tells them something new: the debate is occurring at the wrong altitude. The roughly 40 proposed hypotheses are not 40 competing answers to the same question. They are many answers to a purpose question built on top of one well-compelled answer to the behavior question. The proliferation that looks like C3-diagnostic chaos is actually C2b-diagnostic convergence viewed from the wrong level of description.

"Stop debating which behavior is correct — they're all the same behavior — and recognize that the argument is about purpose" is a genuine intervention in the archaeological literature on this object. It does not resolve the purpose question (the object cannot do that), but it correctly locates where resolution has already occurred and where indeterminacy genuinely remains.

8.4 Material Uncertainty and the Abstraction Failure

Claude's self-diagnosis is the most precise: "I let material uncertainty propagate upward into behavioral uncertainty." When the material passing through the perforation is unknown, the systems treated that uncertainty as uncertainty about the interaction itself — when in fact the interaction is defined at a level of abstraction that is indifferent to the input material. The constraint field operates on whatever passes through the hole, regardless of what it is.

This is a specific reasoning failure: uncertainty at a lower level (input material) contaminates assessment at a higher level (compelled interaction), rather than the system recognizing that the higher-level description is stable across the lower-level variations. The framework's contribution is precisely the insistence that you identify the level at which the constraint field closes, and do not let unresolved variables at other levels prevent you from recognizing closure where it exists.

9. Implications for the Framework

9.1 The Behavior/Purpose Distinction is the Framework's Sharpest Contribution

The experiment identifies the behavior/purpose distinction as the specific point where the framework does work that does not happen naturally — not in the archaeological literature, and not in the default reasoning of any major language-model system. Every other element of the framework (intake screens, evaluation gates, constraint extraction, the worked failure template) was applied correctly by all systems without prompting. The altitude test is the one move that requires explicit demonstration.

9.2 The Distinction Needs a Worked Example to Land

Four systems with access to the framework document failed to apply the behavior/purpose distinction correctly on a live case. This suggests that the distinction as currently stated is not self-teaching — a reader can understand the definitions in the abstract without successfully applying them when faced with an object where material-input variation mimics behavioral variation. A worked example — specifically, the bâton percé case — makes the distinction visible before the reader encounters a live classification.

9.3 The Bâton Percé as Calibration Instrument

The bâton percé functions as a calibration case for the framework itself. Anyone who classifies it as C3 has not internalized the behavior/purpose distinction. Anyone who classifies it as C2b has. This makes it a natural companion to the dodecahedron in any methodological exposition: the dodecahedron demonstrates what C3 looks like when behavior-level indeterminacy is genuine; the bâton percé demonstrates what C2b looks like when it is disguised as C3 by purpose-level proliferation. Together, the two objects make the behavior/purpose distinction visible from both sides.

10. Implications for AI Reasoning

10.1 Context Conditioning vs. Independent Derivation

If the framework's most important distinction requires prior demonstration before language models apply it correctly, then convergence across systems after calibration does not demonstrate independent derivation. It demonstrates that four systems can follow a worked example. This is useful — it confirms the distinction is learnable and uncontroversial — but it is a weaker claim than "four systems independently derive the framework's classifications." The defensible claim is narrower: the systems reliably reach C3 for the dodecahedron even uncorrected (because the dodecahedron's behavior-level indeterminacy is genuine), but they reliably distinguish C2b from C3 only when the altitude test has been made explicit.

10.2 Gravitational Pull of Training Data

The uniform failure across systems suggests that the altitude error reflects a gravitational pull toward the conventional archaeological framing, where every proposed "use" gets counted as a separate hypothesis regardless of what level of description it operates at. This framing dominates the training corpora for all four systems. The framework's contribution is precisely the insistence that you sort by level before you count — a move that cuts against how the source literature organizes the information.

This has implications for any domain where language models are used to reason about classification or categorization: default reasoning will reproduce the organizational structure of the training data, even when a different organizational principle would produce more accurate results.

11. Experimental Limitations

The sample size is small: four systems, one uncorrected run, one corrected run, and one confirmation run each. The result is uniform, which strengthens the finding, but replication with varied prompts and varied test objects would increase confidence.

The corrective prompt is specific and directive — it states the distinction explicitly rather than pointing in a direction. A weaker intervention would better test whether the systems can complete the reasoning once nudged.

The order is not counterbalanced. Every system received the uncorrected prompt first, the corrected prompt second, and the dodecahedron third. A design in which some systems received the correction as a worked example on a different object before the test case would better isolate the effect of the analytic distinction from the effect of being told the first answer was wrong.

ChatGPT's first-attempt failure to identify the artifact introduces a confound for that system — it had already received more guidance than the other three before reaching the altitude error.

Despite these limitations, the uniformity of the result — same error, same correction, zero resistance — is robust enough to support the core finding.

12. Conclusion

The bâton percé paired-run experiment demonstrates four things.

First, the framework's behavior/purpose distinction is the specific point where it does analytic work that does not happen naturally. Every other element of the framework was applied correctly by all systems without prompting.

Second, the distinction, while non-obvious, is uncontroversial. No system resisted the correction. The move from C3 to C2b, once the altitude error is identified, is recognized as obviously correct by every system tested.

Third, the correction does not over-generalize. All four systems, after calibration on the bâton percé, correctly held C3 for the Roman dodecahedron — an object whose behavior-level indeterminacy is genuine. As ChatGPT stated in the most concise formulation produced: "The baton plausibly has one hole-mediated force-transfer behavior with many possible ends; the dodecahedron has many possible behaviors before you even get to the question of ends."

Fourth, the bâton percé functions as a calibration instrument for the framework itself. It catches whether an analyst has internalized the behavior/purpose distinction before they reach harder cases. Paired with the dodecahedron — which confirms C3 through the same altitude test applied in the opposite direction — it makes the framework's most important analytic contribution visible, testable, and teachable.

The complete dataset — twelve classifications across four systems, with uniform error patterns, uniform correction, zero resistance, and zero over-application — provides empirical evidence that the framework's central distinction is both substantive and non-obvious.