Benchmark

90 Metrics, Complete
Transparency

Every claim on this site is backed by reproducible benchmarks. We show our wins, our losses, and even our overfitting analysis.

ColorBench: deterministic, float64, 13 categories. All data, scripts, and checkpoints are open source. No cherry-picking. Full train/test split analysis included.

62 GenSpace wins
19 ties
9 OKLab wins

83 internal + 7 independent validation = 90 total metrics

MetricSpace

Color Difference Accuracy

MetricSpace predicts human color perception more accurately than any existing standard, including the industry-standard CIEDE2000 formula.

STRESS (CIE 217:2016) evaluated on COMBVD (3,813 pairs from 6 sub-datasets) and the held-out MacAdam 1974 set (128 pairs). Lower is better. Our own human-feedback set is reported separately below — rank-order only.

Our Own Human-Feedback Set rank-order only

3,477 judgements by 71 observers over 47 color pairs, collected on helmlab.space. We audited this set and do NOT publish STRESS on it: ratings are 5-level categorical, which breaks interval-scale metrics — any STRESS number from it (ours included) is a scale artifact. Rank order is the honest statistic. Two more caveats: pairs were sampled from helmlab's own hue zones (not independent), and with 47 pairs the MetricSpace–CIEDE2000 gap is only marginally significant (Δρ 95% CI [+0.002, +0.117], paired bootstrap).

MetricSpace
ρ = 0.954
CIEDE2000
ρ = 0.907
CIE Lab ΔE76
ρ = 0.906
OKLab
ρ = 0.885

More Held-Out Datasets

STRESS on independent psychophysical sets MetricSpace was never trained on (same protocol: Bradford CAT to D65, CAM16-UCS via colour-science with a gray-ramp sanity check). One of these we lose — published anyway.

Dataset MetricSpace CIEDE2000 CAM16-UCS CIE Lab OKLab
Munsell renotation (neighbor pairs) n=3590
constant perceptual spacing between adjacent chips; illuminant C, Bradford-adapted
30.34 42.94 43.70 43.99 51.95
He 2022 (wide-gamut display pairs) n=82
10° observer data, DV published in CIELAB-scaled units — we lose this one
35.89 32.58 34.42 30.77 49.04
GenSpace

Generation Benchmark: 62-9

GenSpace is purpose-built for creating colors: gradients, palettes, gamut mapping. Head-to-head against OKLab (the current CSS standard), it wins 62 out of 90 metrics.

83 internal metrics (deterministic, float64) + 7 independent validation metrics. Opponent: OKLab with standard Euclidean deltaE. Same test harness, same precision.

62
GenSpace wins
9
OKLab wins
19
Ties
13
Categories
By Category

Category Breakdown

Click any card to expand and see every individual metric in that category.

Performance breakdown across 13 test categories. Click any card to see individual metrics.

83 internal metrics + 7 independent validation metrics in 13 categories. All deterministic, float64.

Gamut

24W3T0L

How well the space maps to real device screens

Cusp validity, boundary smoothness, clipping across sRGB/P3/Rec2020

27 metrics

Application

8W3T1L

Real-world tasks like palettes, tints, and accessibility

Palette generation, gamut mapping, WCAG contrast, animation

12 metrics

Independent

6W0T1L

Tests using data we never trained on

Hung-Berns, Ebner-Fairchild, Pointer's Gamut

7 metrics

Gradient

5W3T3L

How smooth colors blend between two endpoints

CV of perceptual step size, hue drift, banding metrics

11 metrics

Perceptual

5W0T0L

Agreement with how humans actually see color

Munsell, MacAdam, Hung-Berns hue linearity validation

5 metrics

Structural

4W2T2L

Mathematical properties that affect reliability

Hue reversals, OOG excursion, chroma amplification, LMS

8 metrics

Achromatic

2W0T0L

Perfect grays without color contamination

Gray ramp chroma residual under sRGB and D65

2 metrics

Advanced

2W4T0L

Edge cases and stress tests

1000-trip roundtrip, Jacobian condition, 8-bit precision

6 metrics

Hue

2W0T0L

Whether hue labels match human expectation

Hue RMS vs Munsell, primary lightness range

2 metrics

Special

2W0T1L

Problem areas where OKLab is known to struggle

Yellow chroma, blue-to-white midpoint, red-to-white shift

3 metrics

Accessibility

1W0T1L

Usability for colorblind viewers

CVD simulation minimum step deltaE (protan/deutan)

2 metrics

Banding

1W1T0L

Visible stepping artifacts in gradients

Invisible step ratio, duplicate 8-bit bucket count

2 metrics

Numerical

0W3T0L

Mathematical precision of conversions

Round-trip error across sRGB, P3, Rec2020 (float64)

3 metrics
Full Data

Metric Explorer

Search and filter all 90 benchmark metrics. Every number is reproducible.

Sortable, filterable table of all metrics. Values are from the canonical ColorBench run: GenSpace v0.11.1 vs OKLab, both at float64 precision.

MetricCategoryOKLabGenSpaceWinner
CVD deutan min step dEΔEAccessibility0.160.11OKLab
CVD protan min step dEΔEAccessibility0.130.13GenSpace
Gray ramp pure C*C*Achromatic7.61e-71.88e-15GenSpace
Gray ramp sRGB C*C*Achromatic6.04e-71.76e-15GenSpace
1000-trip RTmax ΔEAdvanced6.17e-132.53e-14GenSpace
8-bit exact/10KcountAdvanced10,00010,000Tie
Animation frame-to-frame CV%Advanced55.353.4GenSpace
Channel mono violationscountAdvanced00Tie
Cross-gamut amplification×AdvancedTie
Jacobian conditionAdvanced6.496.47Tie
Chroma preservation (no mud)Application0.4140.41Tie
Data viz min pairwise dEΔEApplication14.0813.27OKLab
Eased animation CV%Application55.155.3Tie
Muddy gradients (C drop >50%)countApplication1212Tie
Multi-stop gradient CV%Application37.436.9GenSpace
Palette harmony accuracy°Application11.79.1GenSpace
Palette L* spacing%Application78.976.5GenSpace
Photo gamut map fidelity°Application0.980.96GenSpace
Shade palette hue drift°Application8.66GenSpace
Shade palette worst hue drift°Application20.920.4GenSpace
Tint/shade hue preservation°Application8.87.9GenSpace
WCAG midpoint contrast:1Application2.732.88GenSpace
Duplicate 8-bit steps%Banding16.113.8GenSpace
Invisible gradient steps%Banding96.797Tie
Cusp smoothness (max jump)Gamut0.8050.075GenSpace
Gamut volume fill%Gamut11Tie
P3 boundary bad huescountGamut1214GenSpace
P3 boundary continuityGamut0.4440.079GenSpace
P3 boundary mean jumpGamut0.020.003GenSpace
P3 cliff max%Gamut0.20.2Tie
P3 cusp mean smoothnessGamut0.0080.005GenSpace
P3 cusp smoothnessGamut0.7780.043GenSpace
P3 invalid cuspscountGamut520GenSpace
P3 mono violationscountGamut710GenSpace
P3 valid cuspscuspsGamut308/360360/360GenSpace
Rec2020 boundary bad huescountGamut13020GenSpace
Rec2020 boundary continuityGamut0.5620.248GenSpace
Rec2020 boundary mean jumpGamut0.0250.006GenSpace
Rec2020 cliff max%Gamut0.70.2GenSpace
Rec2020 cusp mean smoothnessGamut0.0070.006GenSpace
Rec2020 cusp smoothnessGamut0.7560.157GenSpace
Rec2020 mono violationscountGamut601GenSpace
Rec2020 valid cuspscuspsGamut360/360360/360Tie
sRGB boundary bad huescountGamut12315GenSpace
sRGB boundary continuityGamut0.5450.301GenSpace
sRGB boundary mean jumpGamut0.020.005GenSpace
sRGB cliff max%Gamut0.70.2GenSpace
sRGB cusp mean smoothnessGamut0.0090.005GenSpace
sRGB invalid cuspscountGamut610GenSpace
sRGB mono violationscountGamut880GenSpace
sRGB valid cuspscuspsGamut299/360360/360GenSpace
3-color gradient CV%Gradient26.5923.42GenSpace
Banding meanGradient1.81.8Tie
Bright gradient CV (L>0.6)%Gradient23.5225.43OKLab
Cross-lightness gradient CV%Gradient15.5111.21GenSpace
Dark gradient CV (L<0.4)%Gradient46.5333.68GenSpace
Gradient CV (mean)%Gradient30.2730.12Tie
Gradient CV (p95)%Gradient138.2138.73Tie
High-chroma gradient CV%Gradient21.7619.27GenSpace
Max hue drift (non-crossing)°Gradient112.777.5GenSpace
Near-achromatic gradient CV%Gradient79.26102.1OKLab
Worst-case gradient CV%Gradient362.2500OKLab
Hue RMS°Hue30.127.5GenSpace
Primary L rangeHue0.5160.6GenSpace
Ebner-Fairchild hue surfaces (max)°Independent8.18.6OKLab
Ebner-Fairchild hue surfaces (mean)°Independent2.232.11GenSpace
Hung-Berns hue linearity (max)°Independent25.525.2GenSpace
Hung-Berns hue linearity (mean)°Independent4.964.72GenSpace
Pointer gamut boundary smoothnessIndependent0.1320.125GenSpace
Pointer gamut chroma isotropyIndependent0.4130.404GenSpace
Pointer gamut hue uniformityIndependent0.370.261GenSpace
Round-trip P3 16.7Mmax ΔENumerical1.67e-151.89e-15Tie
Round-trip Rec2020 2.1Mmax ΔENumerical1.55e-151.67e-15Tie
Round-trip sRGB 16.7Mmax ΔENumerical1.78e-151.78e-15Tie
Hue agreement with CIE Lab°Perceptual8.58.3GenSpace
Hue leaf constancy°Perceptual73.359.8GenSpace
MacAdam isotropyratioPerceptual1.991.78GenSpace
Munsell Hue spacing%Perceptual18.511.4GenSpace
Munsell Value uniformity%Perceptual2.80.16GenSpace
Blue-White midpoint G/RSpecial1.4081.514GenSpace
Red-White midpoint G-BSpecial0.0620.063OKLab
Yellow chromaSpecial0.2110.333GenSpace
Extreme chroma amplification×Structural5.79×3.79×GenSpace
Hue reversal max angle°Structural30.6GenSpace
Hue reversals (count)countStructural8067GenSpace
Negative LMS colors%Structural00Tie
OOG excursion pairs%Structural9.89.8Tie
OOG max distanceStructural0.110.103GenSpace
Primary hue disc (P3)°Structural1.081.37OKLab
Primary hue disc (sRGB)°Structural1.311.65OKLab
Showing 90 of 90 metrics
62 GenSpace19 Ties9 OKLab
Independent Validation

Tested on Data We Never Trained On

Three independent datasets from published color science research (1980-1998). GenSpace wins 6-1 against OKLab on data it never saw.

Hung & Berns 1995 (hue linearity, 168 samples), Ebner & Fairchild 1998 (constant-hue surfaces, 321 samples), Pointer 1980 (real surface color gamut, 576 boundary points). None used in optimization.

Hung & Berns 1995

168 samples

Do straight lines in the color space match straight lines in human hue perception?

Hue linearity: angular deviation from constant-hue lines. 12 hues, 13 targets each, 9 observers.

Red
2.42 vs 2.84
Red-yellow
4.24 vs 3.55
Yellow
5.71 vs 4.51
Yellow-green
5.48 vs 4.69
Green
1.64 vs 1.69
Green-cyan
7.13 vs 8.77
Cyan
9.37 vs 9.94
Cyan-blue
7.01 vs 5.58
Blue
4.62 vs 3.29
Blue-magenta
4.52 vs 3.76
Magenta
3.59 vs 3.54
Magenta-red
3.78 vs 4.43
Score 6W 1T 5L

Ebner & Fairchild 1998

321 samples

When you change lightness and chroma but keep the hue name the same, does the color space agree?

Constant perceived-hue surface deviation. 15 hues. Mean and max angular deviation from ideal.

Space Mean Max
CIE Lab 2.95 16.0
OKLab 2.23 8.1
GenSpace 2.10 8.6
GenSpace wins mean deviation. OKLab wins max deviation.

Pointer's Gamut 1980

576 pts

How uniformly does each space represent real-world surface colors?

Real surface color boundary (16 L levels, 36 hue angles). Chroma CV, boundary smoothness, hue uniformity.

Space C* CV Smooth Hue CV
CIE Lab 0.479 0.144 0.034
OKLab 0.413 0.132 0.370
GenSpace 0.404 0.125 0.262
CIE Lab wins hue uniformity because Pointer's data is defined in CIE Lab coordinates.

Independent Validation Total

Across 3 published datasets (1980-1998), none used in training

6 - 1 (7 metrics)
Honesty Check

Overfitting Analysis

We optimized MetricSpace on color difference data. Could it have just memorized the answers? We tested this honestly and show you the results.

80/20 stratified split (seed=42), multiple DOF configurations. Train-test gap exists (+1.8) but held-out test still beats all competitors.

Does the model genuinely predict color perception, or did it just memorize the training data? We tested this rigorously with held-out data the model never saw during training.

80/20 train-test split (seed=42, 3050/763 pairs). Multiple DOF configurations tested. Cross-validated estimate: STRESS 24.3.

ModelParamsDOFTrainTestGap
v20b baseline027.7227.57-0.15
v21 (full-data)7222.1423.91+1.77
Phase 1 train-only625.3525.65+0.30
Phase 1+2 train-only4822.7824.59+1.82

Key Findings

Mild overfitting exists. +1.8 STRESS gap between train and held-out test. This is real and acknowledged.
Gap is from data variance, not memorization. v21 (full-data) gap = +1.77, train-only gap = +1.82. Nearly identical.
Held-out test still beats all competitors. Held-out test STRESS (24.59) vs CIEDE2000 (29.20) on full dataset = 16% better.
MacAdam generalizes independently. MacAdam 1974 (never trained on) = 19.12, consistent with v21's 19.51.
Low DOF shows near-zero gap. 6 DOF shows +0.30 gap. Overfitting scales with parameters but stays controlled.

Published STRESS: 22.48 (full-data optimized) | Cross-validated: ~24.3 | Both are still #1 among all tested competitors.

MetricSpace

When to Use MetricSpace

MetricSpace is purpose-built for color difference prediction — not generation. Use it when you need to measure, not create.

Quality Control

Print, display, textile color matching. 23% lower STRESS than CIEDE2000 on COMBVD.

Color Matching Tolerance

Pair-dependent SL/SC weighting adapts to the specific lightness and chroma of each color pair.

A/B Testing

Ranks real user judgments better than CIEDE2000 (Spearman 0.954 vs 0.907) on our own 47-pair human-feedback set — rank-order only, see note below.

Accessibility Checking

Euclidean deltaE that's actually perceptually calibrated. OKLab STRESS = 47.35 — not designed for distance prediction.

Research

Transparent pipeline, fully invertible, open source. All parameters, datasets, and optimization scripts published.

Verdict

Are We the Best?

Color difference measurement: Yes.

MetricSpace v21 achieves the lowest published STRESS on COMBVD, is within 0.8 of the best (CAM16-UCS) on the held-out 128-pair MacAdam set, and ranks our own human-feedback set better than CIEDE2000 (Spearman 0.954 vs 0.907). No other metric matches human perception this accurately across this range of datasets. Caveat: cross-validated estimate is ~24.3 (not the published 22.48), and CIEDE2000 wins on 3 of 6 COMBVD sub-datasets.

Generation tasks: Best-rounded, not best at everything.

GenSpace wins 62-9 vs OKLab across 90 metrics, including 6-1 on independent 3rd-party datasets OKLab was optimized on. However: OKLab is better for near-achromatic gradients (22%), CVD deutan palettes (43%), and native CSS oklch(). CIE Lab's hue angles remain the established industry reference for hue naming.

Overall: First to do both.

Helmlab is the first color space library to achieve state-of-the-art in both perceptual color difference measurement and visual generation quality simultaneously. No other space does both.

Methodology

How We Test

Deterministic

Every metric is computed at float64 precision with fixed seeds. Run the same code, get the same numbers. No stochastic variation.

Head-to-head

Same test harness, same input colors, same precision for both spaces. Winner is determined by the metric's natural direction (lower or higher is better).

No Cherry-picking

All metrics are reported, including our 9 losses. We do not add or remove metrics based on whether we win them.

Open Source

ColorBench source code, all data files, and checkpoint parameters are publicly available on GitHub for independent verification.

What We Did NOT Test

  • HDR color differences — no HDR psychophysical dataset available
  • Cross-surround conditions — all data is standard viewing conditions
  • Display-specific gamuts — only standard sRGB / Display P3 / Rec.2020 primaries
  • Computational performance — not formally benchmarked (JS microbenchmark: GenSpace ≈2.5–2.7× OKLab per color)
  • Perceptual ranking with human observers — GenSpace metrics test geometric/mathematical properties, not direct human preference