# Inputs and evaluation workflow

## From interview to decision

Bring your recorded interview and a quality transcript. Set up your account and company profile. Add the job, must-haves, interview questions, and ideal answers. Add the candidate and their transcript. Review the comprehensive alignment score, source evidence, and gaps before making the final human decision. A score is available only when the configured evaluation requirements are met.

## Public input contract

A useful evaluation begins with role specific information. The public input contract contains:

1. **Business context:** the delivery objective, relevant team topology, dependencies, and expected outcomes.
2. **Job description:** the scope, responsibilities, and operating environment of the role.
3. **Must haves:** explicit capabilities or experience that are required for this role.
4. **Approved interview questions:** the questions used to gather evidence.
5. **Ideal answer criteria:** observable job criteria that a strong answer should address.
6. **Interview transcript:** the attributable words available for evaluation.
7. **Administrative metadata:** only what is needed to identify the role, interview, source, and authorized review state.
8. **Opening baseline:** when interview-integrity review is requested, an open-ended work and career-experience question, with its source timing and speaker attribution preserved.

Weak, incomplete, or conflicting inputs reduce what the report can responsibly establish. The system should show that limitation instead of hiding it.

## Public workflow

### 0. Establish an opening baseline when integrity review is requested

Start with an open question about the candidate’s path into engineering, work experience, and why the role fits. This gives the candidate room to explain their own experience before harder technical prompts and creates a comparison sample for later source review. It is not a personality, psychological, calm, or capability score.

TeamStation AI research treats the opening as a possible low-pressure primer that may help a candidate settle into the interview. That is a company-reported design observation requiring separate validation. The question must stay focused on work and career experience, and should not solicit protected or unnecessary personal information.

Only a usable baseline can support a baseline-conditioned speech-pattern review. If the opening answer is missing, interrupted, heavily prompted, or not preserved in the recording, the reviewer should not make that comparison. Pauses, slower speech, accent, disfluency, code-switching, translation, disability, and L2 or ESL communication require context and are not proof of AI use.

### 1. Bind the role context

The evaluation begins with the configured role criteria. The role definition controls relevance. A general impression of technical ability does not replace the requirements of the specific job.

### 2. Establish the question map

Each approved question is connected to its ideal answer criteria. Ideal answer criteria are comparison references, not a script the candidate must repeat word for word.

### 3. Establish answer opportunity

The review identifies whether the candidate had a valid opportunity to address the criterion. A criterion should not be treated as disproven merely because the interviewer never asked about it, interrupted the answer, or supplied an unusable transcript segment.

### 4. Locate attributable evidence

Candidate statements are connected to the relevant question and criterion. Interviewer statements, third party claims, and unattributed transcript text should not be presented as the candidate's evidence.

### 5. Record support, gaps, and contradictions

The report distinguishes evidence that supports a criterion, evidence that supports only part of it, evidence that explicitly conflicts with it, and evidence that was not observed. A gap is not automatically a contradiction.

### 6. Summarize observable work reasoning

The system organizes demonstrated job related reasoning, such as problem decomposition, decision explanation, ownership, adaptation, and validation. The summary remains bounded to the interview evidence.

### 7. Apply configured processing

After evidence is established, the product applies its configured processing. The formulas, weights, thresholds, score anchors, calibration logic, and internal controls are proprietary and are not published in this knowledge layer.

In the governed product contract, this numeric processing is performed by versioned software after the evidence state is locked. The semantic evidence stage is not permitted to select scores, weights, anchors, gates, recommendations, or release decisions.

### 8. Open the human review gate

A qualified reviewer inspects the role configuration, transcript quality, evidence connections, gaps, contradictions, and any output used in an employment decision. The reviewer can correct, exclude, or request more evidence.


## From interview answers to role alignment

Axiom Cortex connects what the candidate said, how the answer addresses the problem, and what the role requires. The report brings the alignment result, supporting evidence, critical gaps, and follow-up questions into one review.

### How an answer becomes a finding

1. **Lock the role reference:** Bind the job, must-haves, exact questions, and approved ideal-answer criteria before reviewing candidate evidence.
2. **Build each Answer Evaluation Unit:** Connect the question, ideal answer, complete attributable answer, and relevant follow-ups. Preserve the original candidate words and source location.
3. **Separate evidence from context:** Identify demonstrated criteria, partial support, contradictions, and gaps. Interviewer suggestions cannot earn candidate credit. Cross-answer context remains traceable to its original question.
4. **Examine the behavioral axioms:** Review accuracy, mental model, procedural knowledge, clarity, and cognitive load against job-related answer evidence.
5. **Calculate the configured result:** Versioned software applies the selected scoring policy to accepted evidence and approved measurements. Role importance, interview coverage, and critical requirements remain visible.
6. **Review alignment and unresolved gaps:** Read the result alongside the Evidence Locker, must-have coverage, cross-answer consistency, and targeted follow-up questions. The hiring team makes the final decision.

### What the mathematics tells the reviewer

The areas below describe documented methods and their intended interpretation. Each evaluation identifies which configured methods actually computed a result. Research methods require their own data, implementations, and validation; inclusion here does not establish production execution.

#### Meaning and conceptual distance

**Methods:** Conceptual Fidelity, Fréchet semantic distance.

Compare the concepts expressed in the answer with the concepts required by the ideal-answer blueprint. Shows where the technical meaning aligns, diverges, or lacks support. Correct paraphrases can preserve meaning without repeating the blueprint.

Distance methods require approved embeddings, comparison inputs, and a calibrated mapping. Semantic similarity alone does not establish technical correctness.

#### Reasoning structure and connections

**Methods:** Discourse analysis, Optimal Transport, Wasserstein distance.

Compare represented relationships among concepts and steps, including the connection between a claim, its mechanism, and its consequence. Helps distinguish a connected explanation from adjacent technical terms and locate differences from the expected reasoning.

A transport distance describes the supplied representations and cost model. It does not prove a reasoning sequence unless those relationships are explicitly represented.

#### Five behavioral axioms

**Methods:** Accuracy, Mental Model, Procedural Knowledge, Clarity, Cognitive Load.

Examine correctness, causal understanding, execution steps, explanation clarity, and handling of technical complexity within each answer. Shows which part of an answer is supported and which needs more evidence. A correct definition and a demonstrated implementation address different criteria.

Measurements stay tied to the approved role rubric. Cognitive Load is an evidence-bound rubric dimension, not a measurement of brain activity, stress, or health.

#### Depth, adaptation, ownership, and collaboration

**Methods:** Latent Trait Inference, Logical Knowledge Depth, Problem-Solving Trajectory Analysis, Context Setting, Metacognitive Calibration.

Connect answer-level findings to mechanisms, problem decomposition, changing constraints, clarification, contribution, stakeholder impact, and acknowledgment of unknowns. Builds the role-specific mental-shape view: architectural instinct, problem-solving agility, learning orientation, and collaborative mindset.

Trait synthesis needs an approved mapping from answer evidence. The separate six-domain human-task-agent research model must not be substituted for a production scoring policy.

#### Uncertainty and the next useful question

**Methods:** Bayesian evidence update, Expected Information Gain.

Model how additional evidence changes an estimate and which approved follow-up could reduce the remaining uncertainty. Directs attention to the gap that matters before a decision, including evidence that could change the current interpretation.

These documented methods require calibrated priors, question mappings, and measurement-error inputs. Missing inputs cannot become an invented confidence interval.

#### Relationships across demonstrated skills

**Methods:** Gaussian Graphical Models, Partial correlation.

Study associations among measured skill dimensions while accounting for other represented dimensions. Research can examine which skill measurements connect and where the available evidence is fragmented.

Requires an adequate dataset and approved model. The V4 specification keeps this in offline validation or shadow research unless those requirements are met; association does not establish causation.

#### Language calibration, reliability, and bias analysis

**Methods:** Translation invariance, Inter-rater reliability, Generalizability analysis, Expected Calibration Error, Differential Item Functioning.

Test whether equivalent meaning receives consistent treatment, whether evaluators agree, and whether measurements vary across questions or relevant study groups. Examines the quality of the measurement itself, including unwanted sensitivity to language form and subgroup differences.

Reliability, calibration, and fairness statistics need suitable study data. Accent, pauses, pronunciation, and protected traits cannot stand in for job-related evidence.

#### Role alignment and critical requirements

**Methods:** Configured aggregation, Critical-requirement gates, Human-task-agent alignment distance.

Combine approved measurements according to the role and preserve the effect of must-have criteria, assessment coverage, and unresolved requirements. Shows the supported fit to a particular delivery role and the gaps an overall average could otherwise conceal.

One complete versioned policy governs a result. The published six-domain distance model studies human-task-agent alignment separately from production aggregation.

### Reading the result

- Read the score together with its question-level evidence, coverage, and critical requirements.
- A question that was not asked creates an assessment gap. It cannot silently become a candidate failure.
- A complete answer that misses a tested criterion differs from missing, truncated, or unattributable evidence.
- A later answer can clarify or conflict with an earlier one; its source and scoring use must remain explicit.
- A language-pattern irregularity is a reason to inspect evidence and ask a targeted follow-up. It does not establish that a candidate used AI or was dishonest.
- A mathematical result describes its evidence and configuration. The hiring team reviews the unresolved gaps and makes the decision.

For the research papers and their study designs, see [Behavioral axioms and the research model](https://axiomcx.dev/knowledge/work-reasoning-dimensions.md).

## Important distinctions

- An ideal answer is a source of observable criteria, not a secret answer key.
- A missing answer is not the same as an incorrect answer.
- Team activity is not automatically personal ownership.
- A confident explanation is not automatically strong evidence.
- A short answer is not automatically weak if it directly supports the criterion.
- A long answer is not automatically strong if it lacks relevant evidence.
- A transcript can support a narrow finding without supporting a broad conclusion.

## Related resources

- [Evidence model](https://axiomcx.dev/knowledge/evidence-model.md)
- [Complete processing review](https://axiomcx.dev/processing-engine/)
- [Observable work reasoning](https://axiomcx.dev/knowledge/work-reasoning-dimensions.md)
- [Human review](https://axiomcx.dev/knowledge/human-review.md)
- [Evidence model JSON](https://axiomcx.dev/data/evidence-model.json)
