Too much, too soon
Correct information is buried under detail the user did not need yet.
UX-centered evaluation for language-model experiences
UXSWE evaluates what happens around the answer—whether an AI experience is economical, well-paced, continuous, and cognitively clear after outcome quality is checked.
Five layersContext to diagnosis
Four pillarsA profile, not a rank
Separately visibleOutcomes and safeguards
Profiles preserve tradeoffs that a single leaderboard number can hide.
The evaluation gap
Traditional quality checks answer whether a result is right. They do not fully explain whether the experience helped a person understand, decide, recover, and finish.
Outcome
The facts are correct.
Experience
The task still stalls.
Correct information is buried under detail the user did not need yet.
The response is accurate while leaving the user without a usable next step.
The answer lacks hierarchy, labels, or grouping that supports scanning.
Errors are surfaced without recovery paths, revisions, or clear alternatives.
The system does not signal status, progress, or when useful work is ready.
Constraints and decisions disappear, forcing the user to rebuild context.
The UXSWE framework
Context determines what good means. Outcome gates establish whether the result is acceptable. Pillars describe the experience. Safeguards and UX laws keep critical risks and diagnoses visible.
Layer 1
Users, goals, tasks, environments, devices, languages, and stakes.
Layer 2
Task success, correctness, calibration, safety, and instruction adherence.
Layer 3
The shape and quality of the experience after outcomes are checked.
Layer 4
Accessibility, transparency, and control remain separately visible.
Layer 5
Explanatory lenses for investigating causes—not an extra score pile.
Evaluation boundary
Correctness, task success, calibration, safety, and instruction adherence stay outside pillar scores. Passing an outcome gate does not prove the experience was good.
Cross-cutting safeguards
Four geometric pillars
The pillars are descriptive lenses, not outcome checks. Read them together as a profile to expose strengths, weaknesses, and tradeoffs.
Tempo is intentionally marked as less empirically calibrated than the other pillars. The framework makes that uncertainty visible.
01
The right amount of relevant information for the user’s moment.
Rewards useful signal, progressive disclosure, and proportionate detail. It does not reward shortness for its own sake.
Boundary: Not correctness and not raw response length.
02
How quickly the interaction produces useful, accepted progress.
Looks beyond first-token latency to time-to-first-useful-progress, feedback cadence, and time-to-accepted-result.
Boundary: Not infrastructure speed or TTFT alone.
03
The system’s ability to sustain coherent progress across turns.
Measures context retention, recoverability, revision quality, and whether the interaction reaches a usable completion.
Boundary: Not a proxy for prompt length or memory capacity.
04
How clearly the response organizes meaning for comprehension and action.
Evaluates hierarchy, grouping, findability, scanability, and whether the organization matches the user’s mental task.
Boundary: Not visual polish or formatting volume.
Benchmarking methodology
A benchmark should preserve interaction traces, distinguish deterministic gates from judged experience, and disclose the uncertainty behind every comparison.
Specify users, intended outcomes, environments, constraints, stakes, and representative task strata.
Use controlled repeats, counterbalancing where needed, and enough coverage to expose variance—not one-off demos.
Preserve prompts, outputs, revisions, events, timestamps, tool calls, and acceptance or abandonment signals.
Evaluate task success, correctness, calibration, safety, and instruction adherence deterministically where possible.
Combine validated instrumentation, human ratings, and model-assisted judgment with explicit rubrics and audit samples.
Publish distributions, confidence, failure strata, judge agreement, and pillar profiles before any composite score.
Metrics and benchmark preview
A composite can support a decision, but it should never be the first or only view. UXSWE keeps outcome gates, safeguards, pillar shape, variance, and uncertainty inspectable.
Illustrative schema example — not measured benchmark results.
Every value below demonstrates a reporting structure derived from the documentation schema.
Illustrative pillar profile
Scale · 0–100Economy
Tempo
Continuity
Structure
Illustrative outcome gates
Outcome quality
UX mean
UX balance
Effective UXSWE
Schema illustration: Information Economy 78, Interaction Tempo 91, Interaction Continuity 69, Cognitive Structure 84. Tempo is strongest; continuity remains the limiting experience dimension. A total alone would obscure that.
Research and evidence ledger
UXSWE draws from usability standards, accessibility guidance, psychology-informed UX heuristics, human–AI interaction research, and current model-evaluation work. Each source supports a construct or method—not the framework as a universal standard.
Evidence tiers describe source maturity, not a numerical confidence score. Standards define requirements; peer-reviewed and validated work supports methods; preprints contribute emerging evidence; practitioner sources remain diagnostic or generative.
Documentation overview
The working documentation is organized as a research and implementation bundle. Public routes are not available yet, so this overview describes what is covered without pointing to placeholder pages.
The five-layer model, evaluation boundaries, and how profiles replace a misleading single number.
Definitions, example signals, observable measures, and safeguards kept outside pillar scores.
Task design, repeated trials, judge protocols, aggregation, uncertainty, and statistical reporting.
Implementation-ready data structures, reference ledger, validation notes, and corpus audit material.
A broader definition of quality