Skip to content

UX-centered evaluation for language-model experiences

Correct is not the same as good to use.

UXSWE evaluates what happens around the answer—whether an AI experience is economical, well-paced, continuous, and cognitively clear after outcome quality is checked.

Five layersContext to diagnosis

Four pillarsA profile, not a rank

Separately visibleOutcomes and safeguards

Evaluation profile
Illustrative
/100
Economy78
/100
Tempo91
/100
Continuity69
/100
Structure84

Profiles preserve tradeoffs that a single leaderboard number can hide.

The evaluation gap

A correct answer can still fail the user.

Traditional quality checks answer whether a result is right. They do not fully explain whether the experience helped a person understand, decide, recover, and finish.

Outcome

The facts are correct.

Experience

The task still stalls.

01Verbosity

Too much, too soon

Correct information is buried under detail the user did not need yet.

02Incompleteness

Right, but unfinished

The response is accurate while leaving the user without a usable next step.

03Poor structure

Hard to inspect

The answer lacks hierarchy, labels, or grouping that supports scanning.

04Weak recovery

Failure becomes a dead end

Errors are surfaced without recovery paths, revisions, or clear alternatives.

05Feedback gaps

Silence feels slow

The system does not signal status, progress, or when useful work is ready.

06Context loss

Every turn starts over

Constraints and decisions disappear, forcing the user to rebuild context.

The UXSWE framework

Evaluate the system in layers.

Context determines what good means. Outcome gates establish whether the result is acceptable. Pillars describe the experience. Safeguards and UX laws keep critical risks and diagnoses visible.

  1. 01

    Layer 1

    Context of use

    Users, goals, tasks, environments, devices, languages, and stakes.

  2. 02

    Layer 2

    Outcome gates

    Task success, correctness, calibration, safety, and instruction adherence.

  3. 03

    Layer 3

    Four pillars

    The shape and quality of the experience after outcomes are checked.

  4. 04

    Layer 4

    Safeguards

    Accessibility, transparency, and control remain separately visible.

  5. 05

    Layer 5

    Diagnostic UX laws

    Explanatory lenses for investigating causes—not an extra score pile.

Evaluation boundary

Correctness, task success, calibration, safety, and instruction adherence stay outside pillar scores. Passing an outcome gate does not prove the experience was good.

Cross-cutting safeguards

  • 01Accessibility & inclusion
  • 02Trust & consequence transparency
  • 03User control & recoverability

Four geometric pillars

A shape for each dimension of use.

The pillars are descriptive lenses, not outcome checks. Read them together as a profile to expose strengths, weaknesses, and tradeoffs.

Tempo is intentionally marked as less empirically calibrated than the other pillars. The framework makes that uncertainty visible.

Converging evidence

01

Information Economy

The right amount of relevant information for the user’s moment.

Rewards useful signal, progressive disclosure, and proportionate detail. It does not reward shortness for its own sake.

  • Useful-density
  • Relevance
  • Progressive disclosure

Boundary: Not correctness and not raw response length.

Moderate evidence · less calibrated

02

Interaction Tempo

How quickly the interaction produces useful, accepted progress.

Looks beyond first-token latency to time-to-first-useful-progress, feedback cadence, and time-to-accepted-result.

  • Useful progress
  • Feedback cadence
  • Accepted-result time

Boundary: Not infrastructure speed or TTFT alone.

Converging evidence

03

Interaction Continuity

The system’s ability to sustain coherent progress across turns.

Measures context retention, recoverability, revision quality, and whether the interaction reaches a usable completion.

  • Context retention
  • Recovery
  • Completion

Boundary: Not a proxy for prompt length or memory capacity.

Established foundations

04

Cognitive Structure

How clearly the response organizes meaning for comprehension and action.

Evaluates hierarchy, grouping, findability, scanability, and whether the organization matches the user’s mental task.

  • Hierarchy
  • Grouping
  • Findability

Boundary: Not visual polish or formatting volume.

Benchmarking methodology

From context to evidence in six steps.

A benchmark should preserve interaction traces, distinguish deterministic gates from judged experience, and disclose the uncertainty behind every comparison.

Hybrid judgment can combine instrumentation, humans, and models—but rubrics, agreement, and audit samples remain part of the result.
  1. 01

    Define context and tasks

    Specify users, intended outcomes, environments, constraints, stakes, and representative task strata.

  2. 02

    Run repeated trials

    Use controlled repeats, counterbalancing where needed, and enough coverage to expose variance—not one-off demos.

  3. 03

    Capture the interaction

    Preserve prompts, outputs, revisions, events, timestamps, tool calls, and acceptance or abandonment signals.

  4. 04

    Apply outcome gates

    Evaluate task success, correctness, calibration, safety, and instruction adherence deterministically where possible.

  5. 05

    Judge safeguards and pillars

    Combine validated instrumentation, human ratings, and model-assisted judgment with explicit rubrics and audit samples.

  6. 06

    Report profiles and uncertainty

    Publish distributions, confidence, failure strata, judge agreement, and pillar profiles before any composite score.

Metrics and benchmark preview

Profiles reveal what totals erase.

A composite can support a decision, but it should never be the first or only view. UXSWE keeps outcome gates, safeguards, pillar shape, variance, and uncertainty inspectable.

Illustrative schema example — not measured benchmark results.

Every value below demonstrates a reporting structure derived from the documentation schema.

Illustrative pillar profile

Scale · 0–100
01

Economy

02

Tempo

03

Continuity

04

Structure

Illustrative outcome gates

  • Task success84%
  • Correctness81%
  • Calibration72%
  • Safety97%
  • Instruction adherence88%

Outcome quality

UX mean

UX balance

Effective UXSWE

Schema illustration: Information Economy 78, Interaction Tempo 91, Interaction Continuity 69, Cognitive Structure 84. Tempo is strongest; continuity remains the limiting experience dimension. A total alone would obscure that.

Research and evidence ledger

Grounded in converging evidence.

UXSWE draws from usability standards, accessibility guidance, psychology-informed UX heuristics, human–AI interaction research, and current model-evaluation work. Each source supports a construct or method—not the framework as a universal standard.

UXSWE is an original synthesis. Inclusion here does not mean a source’s authors endorse it, and newer preprints still require replication and calibration.

Evidence tiers describe source maturity, not a numerical confidence score. Standards define requirements; peer-reviewed and validated work supports methods; preprints contribute emerging evidence; practitioner sources remain diagnostic or generative.

Documentation overview

Built to be inspected, adapted, and challenged.

The working documentation is organized as a research and implementation bundle. Public routes are not available yet, so this overview describes what is covered without pointing to placeholder pages.

01

Framework

The five-layer model, evaluation boundaries, and how profiles replace a misleading single number.

02

Pillars and metrics

Definitions, example signals, observable measures, and safeguards kept outside pillar scores.

03

Methodology and scoring

Task design, repeated trials, judge protocols, aggregation, uncertainty, and statistical reporting.

04

Schema and evidence

Implementation-ready data structures, reference ledger, validation notes, and corpus audit material.

A broader definition of quality

Evaluate the experience, not only the answer.