# Observable work reasoning

## Definition

Observable work reasoning is the job related reasoning a candidate demonstrates in the supplied interview. It is evidence about the answer, in the context of a configured role. It is not a diagnosis, a personality profile, a brain scan, or a claim about the person's private mental state.

In public Axiom Cortex language, mental shape is shorthand for the pattern formed by these supported observations. Neuro-psychometric alignment refers to mapping that evidence pattern to configured role and delivery requirements. Neither term means brain measurement, clinical psychometrics, or access to private mental states.

<!-- shared:research-guide:start -->
## Behavioral axioms and the research model

The science connects answer-level evidence, work-reasoning patterns, and the demands of human-agent delivery.

The 2025 report combines five B-Axiom checks into four role-fit traits. The 2026 study represents mental shape using six work-reasoning domains. These are related research layers, not interchangeable score definitions.

### Behavioral axioms

Accuracy, mental models, procedural knowledge, clarity, and cognitive load, examined answer by answer.

- **Accuracy:** Technical correctness against the question and ideal-answer criteria.
- **Mental Model:** The explanation of mechanisms, dependencies, and cause and effect.
- **Procedural Knowledge:** The steps used to implement, diagnose, test, and recover.
- **Clarity:** Whether the technical explanation communicates the relevant reasoning.
- **Cognitive Load:** How the answer handles interacting constraints and technical complexity.

Answer Evaluation Units keep each answer connected to its question and evidence before trait synthesis.

Source: [B-Axiom method summary](https://teamstation.dev/hire/by-role/ai-engineer). [Working paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5433476).

### Mental-shape domains

Six dimensions describe how the interview demonstrates understanding, judgment, adaptation, collaboration, learning, and self-calibration.

- **Conceptual Fidelity:** Does the explanation preserve the concepts the system actually depends on?
- **Architectural Instinct:** How are boundaries, dependencies, failure paths, and tradeoffs handled?
- **Problem-Solving Agility:** How does the approach change when new evidence changes the problem?
- **Collaborative Mindset:** How are decisions shared, corrections handled, and handoffs made?
- **Learning Orientation:** How is new evidence used to revise an incomplete or outdated model?
- **Metacognitive Conviction:** Does expressed confidence match the evidence, including what remains unknown?

The 2026 model extends the research framing to human-task-agent alignment.

Source: [Six-domain research explanation](https://teamstation.dev/research/articles/human-task-agent-alignment-stress-test). [Working paper](https://ssrn.com/abstract=7256278).

### Role & agent alignment

Measure the distance between demonstrated reasoning and the requirements of the role, team, and agent workflow.

- **Weighted Euclidean distance:** The published outer model compares a normalized six-domain profile with a task requirement profile.
- **Team-specific requirements:** Separate requirement profiles model stream-aligned, platform, enabling, and complicated-subsystem teams.
- **Agent autonomy:** Autonomy-dependent shifts test how the required human reasoning changes as agents take on more work.
- **Ideal-answer alignment:** In the product, configured role criteria connect the candidate’s answer to the expected technical evidence.

The published outer calculation is a research model; it does not specify the production scoring configuration.

Source: [Human-Task-Agent Alignment working paper](https://ssrn.com/abstract=7256278).

### Meaning & language calibration

Separate technical understanding from surface wording, then apply behaviorally anchored measurement and configured aggregation.

- **Conceptual Fidelity:** Compare the meaning and technical substance of an answer, rather than a memorized phrase.
- **ESL / L2 calibration:** The report describes language-aware calibration to reduce second-language effects on technical evaluation.
- **Behaviorally anchored scoring:** Relate a measurement to observable answer evidence and defined evaluation anchors.
- **Deterministic aggregation:** The research specifies mathematical aggregation from measured signals into higher-level findings.

Source: [Scientific R&D working paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5433476).

### Reliability & bias analysis

Study agreement, error, calibration, subgroup differences, and drift so the measurement itself can be tested.

- **Inter-rater reliability:** Test whether evaluators agree when reviewing the same evidence.
- **ROC and precision-recall analysis:** Examine discrimination, false positives, and false negatives against defined reference outcomes.
- **Calibration error:** Test whether estimated confidence corresponds to observed results.
- **Subgroup fairness:** The report specifies adverse-impact ratios and equal-opportunity gaps for fairness audits.
- **Drift and oversight:** Monitor changes in evaluation behavior, with thresholds and human oversight.

These are documented validation methods, not a claim that every metric has a published production result.

Source: [Scientific R&D working paper](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5433476).

### Delivery physics & sensitivity

Stress-test missing evidence, measurement noise, weighting, review queues, and topology assumptions around the work.

- **Noise and missing domains:** Synthetic perturbations test how uncertain or absent inputs alter the alignment result.
- **Weight sensitivity:** Vary weights to test whether the preferred team topology changes.
- **Kingman queue model:** Study how utilization and variability affect waiting time in the delivery system.
- **Topology-health scenarios:** Explore system-level conditions around human-agent work through Teamlemetry scenarios.
- **Logistic coefficient recovery:** Check whether the analysis recovers effects deliberately planted in synthetic data.

The fixed-seed study uses 24,000 synthetic profiles. It tests the calculation and its assumptions, not real hiring outcomes.

Source: [Human-Task-Agent Alignment working paper](https://ssrn.com/abstract=7256278).

### Working papers and applications

- [AxiomCortex: Scientific R&D Report](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5433476) — foundational working paper · ssrn. Defines Answer Evaluation Units, B-Axiom checks, trait synthesis, Conceptual Fidelity, L2-aware calibration, aggregation, reliability, fairness, and monitoring.
- [Human-Task-Agent Alignment Across Software Team Topologies](https://ssrn.com/abstract=7256278) — working paper · synthetic study. Tests six reasoning domains, weighted alignment distance, team topology, agent autonomy, measurement sensitivity, queue pressure, and synthetic coefficient recovery.
- [Axiom Cortex for LATAM Agentic Engineering](https://teamstation.dev/research/articles/axiom-cortex-latin-america-agentic-engineering-alignment) — company article · teamstation ai. Explains how interview evidence connects to role fit, engineering loops, governance, and delivery alignment.
- [CTO Guide to Agentic Workflow Fit Signals](https://teamstation.dev/research/articles/how-ctos-can-align-the-right-mental-shape-in-their-agentic-ai-dev-workflows) — company article · teamstation ai. Connects work-reasoning signals, team topology, and human-plus-agent engineering workflows.
- [Telemetry Predicts Team Performance](https://teamstation.dev/research/articles/how-telemetry-finds-the-right-mental-shape-and-predicts-team-performance) — company article · teamstation ai. Explores how post-hire delivery signals can test whether interview evidence stays aligned with real work.

### Study design

The 2025 report documents the evaluation framework. The 2026 study stress-tests an outer alignment model on synthetic data. Both are working papers, not peer-reviewed validation. Company articles are not independent scientific validation.

The report supports a human decision, with source evidence, gaps, and follow-up questions available for review.

Source basis: Public SSRN abstracts, the authors’ public method summaries, and the documented Axiom Cortex product contract. Reviewed 2026-09-24.
<!-- shared:research-guide:end -->

## Public observation dimensions

### Problem framing

How the candidate defines the problem, constraints, desired outcome, and relevant operating context.

### Decomposition

How the candidate breaks a larger problem into workable parts, dependencies, stages, or decisions.

### Evidence use

How the candidate identifies useful information, distinguishes observation from assumption, and connects evidence to a conclusion.

### Decision explanation

How the candidate explains a choice, the reason for it, and the conditions that could change it.

### Tradeoff awareness

How the candidate identifies competing goals, costs, risks, or constraints without pretending every option is equally good.

### Validation and feedback

How the candidate checks whether work is correct, observes outcomes, responds to failure, and uses feedback.

### Adaptation

How the candidate updates an approach when requirements, evidence, constraints, or system behavior change.

### Ownership boundaries

How clearly the candidate distinguishes personal actions, shared work, delegated work, and team outcomes.

### Delivery coordination

How the candidate reasons about dependencies, handoffs, sequencing, communication, and impact on the wider delivery chain.

### Risk recognition

How the candidate identifies failure conditions, operational risk, uncertainty, and safeguards relevant to the work.

### Communication clarity

How directly the candidate connects an explanation to the question and role criterion. This dimension concerns the informational content of the answer, not accent, speaking style, disability, fluency, or voice quality.

## Interpretation rules

- A dimension is reported only when the supplied interview contains relevant evidence.
- Evidence from one question may support a narrow observation without supporting a broad profile.
- Not observed means the evidence was not established in the supplied material. It does not mean the person lacks the capability.
- Role relevance controls interpretation. The same answer can have different relevance in different jobs.
- Human review must consider interview design, question quality, transcript quality, and reasonable accommodations.

## Business delivery connection

The dimensions help a reviewer connect demonstrated reasoning to the role's place in a delivery chain. For example, a role with critical cross team dependencies may require stronger evidence about handoffs and risk recognition than a role with a narrow independent scope.

This connection is configured by the company and reviewed by a human. It is not a universal ranking of people.

## Structured resource

See [observable work reasoning dimensions JSON](https://axiomcx.dev/data/work-reasoning-dimensions.json).
