Skip to content

UX-centered evaluation for language-model experiences

From intent to outcome.

Build focused interfaces with UXSWE skills, then evaluate what happens around the answer—whether the experience is economical, well-paced, continuous, and cognitively clear after outcome quality is checked.

Five layersContext to diagnosis

Four pillarsA profile, not a rank

Separately visibleOutcomes and safeguards

Evaluation profile
Illustrative
/100
Economy78
/100
Tempo91
/100
Continuity69
/100
Structure84

Profiles preserve tradeoffs that a single leaderboard number can hide.

The UXSWE skills

Teach agents to build interfaces people can use.

Open-source skills for UI development, UX reasoning, and focused custom interface construction. Build with better instructions, then use UXSWE to evaluate whether the experience actually improved.

Open source

01 / 02

Build the interface, then measure the lift.

Skills → evaluation → improvement

Install the skills

Add the open-source instructions to your agent workflow and start building with a sharper sense of use.

$npx skills add coodier/uxswe-skills

What it covers

01

UI development

Turn product intent into focused, composed interfaces with clear hierarchy and useful defaults.

02

UX reasoning

Give agents a practical vocabulary for information, tempo, continuity, structure, and safeguards.

03

Interaction quality

Make feedback, recovery, user control, and accepted outcomes part of the build—not an afterthought.

The loop

Build, evaluate, improve.

UXSWE is the feedback system for the skills. The first baseline and skills-informed benchmark will be published after the same task set has been measured.

  1. 01

    Build with the skills

    Use the open-source instructions to create a focused interface for a defined task.

  2. 02

    Evaluate the same task

    Run repeated trials and capture the interaction trace with UXSWE gates and pillars.

  3. 03

    Compare the profile

    Report before-and-after shape, safeguards, variance, and uncertainty—not just one score.

  4. 04

    Improve and repeat

    Feed the evidence back into the skills and publish results when the benchmark is measured.

The evaluation gap

A correct answer can still leave the work unfinished.

Traditional quality checks answer whether a result is right. They do not fully explain whether the experience helped a person understand, decide, recover, and finish.

Outcome

The facts are correct.

Experience

The task still stalls.

01Verbosity

Too much, too soon

Correct information is buried under detail the user did not need yet.

02Incompleteness

Right, but unfinished

The response is accurate while leaving the user without a usable next step.

03Poor structure

Hard to inspect

The answer lacks hierarchy, labels, or grouping that supports scanning.

04Weak recovery

Failure becomes a dead end

Errors are surfaced without recovery paths, revisions, or clear alternatives.

05Feedback gaps

Silence feels slow

The system does not signal status, progress, or when useful work is ready.

06Context loss

Every turn starts over

Constraints and decisions disappear, forcing the user to rebuild context.

The UXSWE framework

Map the path from context to outcome.

Context determines what good means. Outcome gates establish whether the result is acceptable. Pillars describe the experience. Safeguards make feedback, recovery, accessibility, transparency, and control visible, while UX laws provide diagnostic vocabulary.

  1. 01

    Layer 1

    Context of use

    Users, goals, tasks, environments, devices, languages, and stakes.

  2. 02

    Layer 2

    Outcome gates

    Task success, correctness, calibration, safety, and instruction adherence.

  3. 03

    Layer 3

    Four pillars

    The shape and quality of the experience after outcomes are checked.

  4. 04

    Layer 4

    Safeguards

    Accessibility, transparency, feedback, recovery, and control remain separately visible.

  5. 05

    Layer 5

    Diagnostic UX laws

    Explanatory lenses for investigating causes—not an extra score pile.

Evaluation boundary

Correctness, task success, calibration, safety, and instruction adherence stay outside pillar scores. Passing an outcome gate does not prove the experience was good.

Cross-cutting safeguards

  • 01Accessibility & inclusionMake the experience perceivable, operable, understandable, and robust across people and contexts.
  • 02Trust & consequence transparencyMake uncertainty, system limits, and meaningful consequences visible when they affect a decision.
  • 03Feedback and recoveryCommunicate state and failure clearly, then offer a straightforward retry, edit, undo, cancel, or safe exit.
  • 04User control & recoverabilityPreserve context and let people steer, revise, pause, or recover without rebuilding the interaction.

Four geometric pillars

Four dimensions of progress.

The pillars are descriptive lenses, not outcome checks. Read them together as a profile to expose strengths, weaknesses, and tradeoffs.

Tempo is intentionally marked as less empirically calibrated than the other pillars. The framework makes that uncertainty visible.

Converging evidence

01

Information Economy

The right amount of relevant information for the user’s moment.

Rewards useful signal, progressive disclosure, and proportionate detail. It does not reward shortness for its own sake.

  • Useful-density
  • Relevance
  • Progressive disclosure

Boundary: Not correctness and not raw response length.

Moderate evidence · less calibrated

02

Interaction Tempo

How quickly the interaction produces useful, accepted progress.

Looks beyond first-token latency to time-to-first-useful-progress, feedback cadence, and time-to-accepted-result.

  • Useful progress
  • Feedback cadence
  • Accepted-result time

Boundary: Not infrastructure speed or TTFT alone.

Converging evidence

03

Interaction Continuity

The system’s ability to sustain coherent progress across turns.

Measures context retention, recoverability, revision quality, feedback continuity, and whether the interaction reaches a usable completion.

  • Context retention
  • Recovery
  • Feedback
  • Completion

Boundary: Not a proxy for prompt length or memory capacity.

Established foundations

04

Cognitive Structure

How clearly the response organizes meaning for comprehension and action.

Evaluates hierarchy, grouping, findability, scanability, and whether the organization matches the user’s mental task.

  • Hierarchy
  • Grouping
  • Findability

Boundary: Not visual polish or formatting volume.

Benchmarking methodology

Follow the evidence from intent to accepted result.

A benchmark should preserve interaction traces, distinguish deterministic gates from judged experience, and disclose the uncertainty behind every comparison. It should also test whether feedback, failure states, and recovery paths help people keep moving.

Hybrid judgment can combine instrumentation, humans, and models—but rubrics, agreement, and audit samples remain part of the result.
  1. 01

    Define context and tasks

    Specify users, intended outcomes, environments, constraints, stakes, and representative task strata.

  2. 02

    Run repeated trials

    Use controlled repeats, counterbalancing where needed, and enough coverage to expose variance—not one-off demos.

  3. 03

    Capture the interaction

    Preserve prompts, outputs, revisions, events, timestamps, tool calls, and acceptance or abandonment signals.

  4. 04

    Apply outcome gates

    Evaluate task success, correctness, calibration, safety, and instruction adherence deterministically where possible.

  5. 05

    Judge safeguards and pillars

    Combine validated instrumentation, human ratings, and model-assisted judgment with explicit rubrics for feedback, recovery, safeguards, and pillars.

  6. 06

    Report profiles and uncertainty

    Publish distributions, confidence, failure strata, judge agreement, and pillar profiles before any composite score.

Metrics and benchmark preview

See the shape of progress, not only a total.

A composite can support a decision, but it should never be the first or only view. UXSWE keeps outcome gates, safeguards, pillar shape, variance, and uncertainty inspectable.

Illustrative schema example — not measured benchmark results.

Every value below demonstrates a reporting structure derived from the documentation schema.

Illustrative pillar profile

Scale · 0–100
01

Economy

02

Tempo

03

Continuity

04

Structure

Illustrative outcome gates

  • Task success84%
  • Correctness81%
  • Calibration72%
  • Safety97%
  • Instruction adherence88%

Outcome quality

UX mean

UX balance

Effective UXSWE

Schema illustration: Information Economy 78, Interaction Tempo 91, Interaction Continuity 69, Cognitive Structure 84. Tempo is strongest; continuity remains the limiting experience dimension. A total alone would obscure that.

Research and evidence ledger

Grounded in converging evidence.

UXSWE draws from usability standards, accessibility guidance, psychology-informed UX heuristics, human–AI interaction research, and current model-evaluation work. Each source supports a construct or method—not the framework as a universal standard.

UXSWE is an original synthesis. Inclusion here does not mean a source’s authors endorse it, and newer preprints still require replication and calibration.

Evidence tiers describe source maturity, not a numerical confidence score. Standards define requirements; peer-reviewed and validated work supports methods; preprints contribute emerging evidence; practitioner sources remain diagnostic or generative.

Documentation overview

Built to be inspected, adapted, and challenged.

The working documentation is organized as a research and implementation bundle. Public routes are not available yet, so this overview describes what is covered without pointing to placeholder pages.

01

Framework

The five-layer model, evaluation boundaries, and how profiles replace a misleading single number.

02

Pillars, safeguards, and metrics

Definitions, example signals, observable measures, and feedback, recovery, and other safeguards kept outside pillar scores.

03

Methodology and scoring

Task design, repeated trials, judge protocols, aggregation, uncertainty, and statistical reporting.

04

Schema and evidence

Implementation-ready data structures, reference ledger, validation notes, and corpus audit material.

A broader definition of quality

Measure the path from intent to outcome.