METHOD HISTORY

How the scoring method changed

Starlight’s scoring methods come from continuing modelling, trial runs and revision. Each stage had reasons that made sense at the time, and limitations discovered later. This page retains both the changes and the unresolved questions.

From whole-person buckets to itemized outcome ledgers

Early methods summarized one person’s performance in six kinds of civilizational contribution. Practice exposed three problems: whether one outcome was counted repeatedly, whether multiple independent outcomes could keep accumulating, and how shared outcomes should be attributed. Later research therefore made the outcome the scoring object and stated facts, value magnitude and personal share separately.

The archive below includes early proposals, the formal ranking baseline and research branches. R numbers identify research lines; RC means a candidate revision. They did not all enter production. A wording change, evidence addition or batch recalculation does not automatically create a new algorithm.

Where the public method currently stands

As of 2026-09-16, Starlight publishes 210 person-level outcome ledgers under one formula and global attribution rule. The remaining 4162 people retain identity, contribution and general introduction, but display none of Canonical v1’s legacy point scores, ranks or T tiers. The site has one current ranking; historical algorithms and results appear only in this archive.

Current outcome ledger
Q = 10M/2 − 1
Lᵢ = aᵢ × Qᵢ
S = 2 × log₁₀(1 + ΣLᵢ)

M is the magnitude of a complete outcome; a is the individual attribution share. A complete outcome first becomes distributable luminance; that luminance is attributed and added, then converted to the displayed score. Both M and a require factual support and assessment reasons. Decimal precision indicates arithmetic precision only.

A key revision: allocating a shared outcome

RC3 placed M×a inside the exponent. For an M=8 outcome credited equally to two people, each received the luminance for M=4: 99 each, or 198 together, while the complete outcome’s luminance was 9,999.

RC3.1 instead computes complete luminance first and gives each person half: 4,999.5 each, still 9,999 together. This changed the accounting rule for shared outcomes, not anyone’s individual treatment. The example shows a mathematical consequence only; whether each person should receive half remains a separate question.

View the current 210-person research ranking →

TIMELINE

Versions and research routes

Listed by archival succession. Early drafts retain date uncertainty where only an archive date exists; research branches may have run in parallel.

2026-09-14 · CURRENT DATA CONVENTION

A common framework for culture, education, arts and science

Reason at the time
To place thought, education, literature, composition and science on one functional scale: knowledge, ethics, aesthetics and teaching need not first turn into technological products, media inventions or randomized-trial outcomes to count as civilizational value.
Problem and limit found
The boundary for a shared M7 remains a contestable editorial judgment. Predecessor shares may consume the original framework twice, and work families may accumulate differently as granularity changes. Independent review still disagrees between M6 and M7 for parts of foundational work; the method does not claim to have eliminated imbalance or stabilized order.
Revised trade-off
Q=10^(M/2)−1, L=aQ and S=2log10(1+ΣL) remain unchanged. One M7 functional gate and shared-budget revisions apply uniformly to affected outcomes. The 210 public ledgers are open to scrutiny; certification is separate.

V0.1 · Early proposal · Early proposal; archived 2026-06-23

Primary and secondary contribution weighting

Reason at the time
To make the most important contribution more visible while still recognizing secondary ones.
Problem and limit found
Neither the primary/secondary distinction nor its weights had an auditable definition. The same contribution could change weight when narrated differently, and the question was entangled with six-bucket aggregation.
What was retained or changed
Contributions should be distinguished by substance, and any weight must have stated reasons.

V0.2 · Early proposal · 2026-01 to 2026-06

Six equal buckets and cross-bucket p=2 power mean

Reason at the time
Equal weights expressed a refusal to pre-rank six kinds of civilizational value; p=2 appeared to keep strong areas visible.
Problem and limit found
A contribution could lose score merely because it was split across more categories. Classification interfered with evaluating the contribution itself.
What was retained or changed
The stance that six values are equal remains; bucket-luminance addition replaced the cross-bucket power mean.

V0.3 · Early proposal · 2026-06-30 to 2026-07-10

Summing bucket luminance

Reason at the time
To separate how bright a bucket is from whether multiple genuinely distinct benefits can accumulate, eliminating the p=2 classification collapse.
Problem and limit found
Ranking, displayed score and star tier still shared one calculation chain, so units and nonlinear transforms affected the threshold for the highest tiers.
What was retained or changed
Luminance addition remains; ranking, displayed score and star tier were separated.

V1.0 · Formal ranking baseline · 2026-07-11 to 2026-09-14

Canonical v1: six buckets, five dimensions, bucket luminance and a three-layer split

Reason at the time
A signed scorecard could be recalculated consistently; six values were equal, strong contributions used nonlinear luminance, and displayed score no longer silently determined a star tier.
Problem and limit found
Judgment still concentrated in whole-person bucket cards. Shared results, overlapping outcomes and accumulation of independent results were difficult to audit item by item.
What was retained or changed
Recalculation, version freezing and the three-layer separation remain; later work made independent outcomes the unit of audit.

R1 · Research branch · 2026-07-12 to 2026-07-15

Separated scoring, ledgering and aggregation

Reason at the time
To place duplicate counting, shared outcomes and aggregation at their proper layers instead of patching them in person cards.
Problem and limit found
The empirical run on gold samples did not pass as a whole. Shares, real relationships and the cost of human review were not closed; correct structure did not make real-person scores usable.
What was retained or changed
Keep the three layers, and record passage at the specification/document level separately from acceptance on real samples.

R1-A · Research branch · 2026-07-25

Repairing U and shares inside the V1.0 whole-person-bucket frame

Reason at the time
It was a small change that could preserve V1.0 readability and existing data.
Problem and limit found
Reducing U and then multiplying a share on the same person bucket double-penalized the same factor. More fundamentally, without an outcome boundary, no share had a stable object.
What was retained or changed
Define outcome boundaries first; do not deduct one factor twice in magnitude and attribution.

R1-B · Research branch · 2026-07-26

Whole-person bucket union repair

Reason at the time
To preserve production continuity as far as possible while preventing one fact from glowing repeatedly through splitting cards, renaming or multiple materials.
Problem and limit found
Treating deduplication as one count per person and category also stopped a second or tenth genuinely independent outcome from accumulating.
What was retained or changed
A genuinely independent second outcome should add something; deduplication cannot become a lifetime bucket cap.

R2 · Research branch · 2026-07-26 to 2026-08-08

Outcome ledger / CACL: only independent results accumulate

Reason at the time
To address two opposite errors directly: claiming the same improvement repeatedly, and failing to accumulate several genuinely independent outcomes.
Problem and limit found
Atomic boundaries in reality, cross-person attribution, a cross-domain common scale and complete evidence closure were unproven. Mechanical conservation did not establish that real people could be scored.
What was retained or changed
Independent outcome ledgers, deduplication and attribution audit entered later research.

R2-A · Research branch · 2026-07-27 to 2026-07-30

Research on a common scale and outcome units

Reason at the time
A genuinely common primitive unit would make comparisons between people and outcomes less dependent on discretion.
Problem and limit found
No validated cross-domain primitive unit was found. Facts, attribution and their correspondence to real cases remained largely unverified, so credible point scores could not follow.
What was retained or changed
A common scale must be tested on cases; mathematical form alone cannot establish objectivity.

R3 · Research branch · 2026-08-10 to 2026-08-11

Epsilon: typed facts and cumulative research

Reason at the time
It separated time, outcome type, deduplication, refinement conservation, attribution and display, and gave unknowns a formal place to stop.
Problem and limit found
Three anonymous axes lacked a reliable correspondence to real quantities. Feeding editorially chosen uniform duration and personal shares in as point values substantially changed tiers and order.
What was retained or changed
Keep typed facts, deduplication and unknown states; stop treating unsupported uniform durations and shares as measurements.

R4 · Research branch · From 2026-08-11

Independent outcome case ledger

Reason at the time
To retain the auditability of an outcome ledger without compressing time and shares lacking factual grounding into precise decimals. Independent outcomes could accumulate, and the object of judgment became concrete and appealable.
Problem and limit found
Case-by-case discretion could still be inconsistent, costly or influenced by fame. An overall grade remains a value judgment, not a natural quantity. Without a public review protocol, blinded review, evidence locations and counterexample tests, it would become arbitrary adjudication; adding many personal credits cannot stand in for total world contribution.
What was retained or changed
Concrete outcomes remain review objects, with counterexamples, reasons and objections retained.

R4.1 · Research branch · From 2026-08-11

Real-person case pilot

Reason at the time
To test whether concrete outcomes from real people could receive reviewable grades.
Problem and limit found
The first two cases mixed unverified status and grade into aggregation. Agreement among similar AI systems could also be shared bias.
What was retained or changed
Keep objections and state separation; do not use vote count as a substitute for evidence.

R4.2 · Research branch · From 2026-08-12

Precedent-anchored calibration

Reason at the time
To compare new cases with documented earlier cases instead of deciding each from zero.
Problem and limit found
Anchors can themselves be wrong; fame, identity leakage and overly broad comparison units can transmit bias to later cases.
What was retained or changed
Anchors must be retractable, and comparisons must state their boundary of application.

R4.3 · Research branch · From 2026-08-12

Complete candidate rule

Reason at the time
To connect outcome admission, deduplication, grading, luminance accumulation and review into a complete rule.
Problem and limit found
A complete written rule is not a validated rule. Review seats, cross-domain lexicons and accumulation curves retained unresolved value choices.
What was retained or changed
Separate mathematical rules, review judgments and real acceptance.

R4.4 · Research branch · From 2026-08-12

Institutional calibration

Reason at the time
To make external review, blinding, sensitivity and public challenge conditions of finalization.
Problem and limit found
A procedure written in advance does not mean qualified reviewers or completed review already exist.
What was retained or changed
Keep evidence of institutional design separate from evidence that the institution has actually run.

R4.5 · Research branch · From 2026-08-12

Value choices and procedural rights

Reason at the time
To show how different curves change order, and to add objection rights for assessed people and co-contributors.
Problem and limit found
Changing a curve can change the relative standing of peak contributions and prolific contributors; information asymmetry and paper-only appeals remain unresolved.
What was retained or changed
A curve cannot be silently decided by a code default; revisions must retain their reasons.

R4.6 · Research branch · From 2026-08-12

Outcome-ladder profile and partial order

Reason at the time
To retain the distribution of outcomes at each level and make a robust comparison only where every level is no lower.
Problem and limit found
It reduces dependence on one curve but cannot yield an intuitive complete one-score ranking for everyone, and it does not cover all synergy or saturation effects.
What was retained or changed
Outcome distributions continue as sensitivity checks; partial order was not adopted as final output for current person pages.

R4.7 · Research branch · From 2026-08-12

Case and governance process

Reason at the time
To organize outcome boundaries, shared contribution, counterfactuals, harms and objections into a reviewable case chain.
Problem and limit found
Structural tests used many synthetic inputs. Real registration, external review and actual operation cannot be replaced by test receipts.
What was retained or changed
Count procedure tests, real cases and external validation separately.

R4.8 · Research branch · From 2026-08-13

V2.0 first replacement candidate

Reason at the time
Internal real-person cases began replaying the accumulation of one and multiple outcomes, attempting to propagate uncertainty through intervals.
Problem and limit found
Interval rankings were not intuitive; outcome granularity, data advantage and curve choice could still alter comparison. Internal AI experiments were not external replication.
What was retained or changed
Keep outcome ledgers and conditional scenarios while continuing to seek reasoned point values and comparable rules.

R4.9 · Research branch · From 2026-08-13

V2.0-RC2: point values and curve calibration

Reason at the time
To restore item-level explanations and attempt explicit rules for point scores and exact ranking.
Problem and limit found
It still depended on five-dimensional synthesis; whole-life ledgers were not delivered. Choosing an aggregation curve before item-level placement was settled amplified error.
What was retained or changed
Explain why each item is assessed as it is before deciding how items accumulate.

R4.10 · Research branch · From 2026-08-14

V2.0-RC3: outcome magnitude and attribution

Reason at the time
To remove five dimensions from the mathematics, use outcome magnitude M and personal attribution a, and include bounded person ledgers and small-contribution scenarios.
Problem and limit found
The candidate applied g=M×a before exponentiation, so it could not allocate the full outcome luminance conservatively. M, a and coverage still needed item-level justification; synthetic lives could not be evidence for ordinary populations.
What was retained or changed
Keep M, a and independent outcomes; RC3.1 later changed the mathematical position of allocation.

V2.0-RC3.1 · Research formula frozen · Specification frozen 2026-08-25

Compute outcome luminance first, then allocate personal shares

Reason at the time
Q=10^(M/2)−1; L=aQ; S=2log10(1+ΣL). An outcome first receives its full luminance, then co-contributors draw shares from one budget; genuinely independent outcomes accumulate.
Problem and limit found
A closed mathematical chain does not establish that magnitudes M, shares a or cross-domain judgment are measured correctly. Formula freezing, case calibration and generalization validation are three different things.
What was retained or changed
Luminance conservation, a global attribution budget, item-level deduplication, and separation of formula from empirical validation became the basis of current research assessment.

V2.0 · assessment protocol · Research assessment adoption · From 2026-09-06

Reasoned point estimates alongside continuing revision

Reason at the time
With the mathematics frozen, facts, magnitude and attribution received stated reasons, yielding contestable point estimates while retaining alternative scenarios and omissions.
Problem and limit found
Richer records do not mean greater contribution. The same retrieval procedure does not assure equally complete lifetime coverage. Point estimates are reproducible but still contain value judgments.
What was retained or changed
Rule revisions recalculate all affected people together; legacy scores and desired ranks do not tune parameters. A twelve-person research batch completed same-standard comparison. The public research ranking later grew to fifty people and expanded to the current two hundred on 2026-09-16.

V2.0 · provisional-evidence rule · Evidence-handling revision · 2026-09-12

Making gaps in hard-to-obtain evidence visible and repairable

Reason at the time
To allow reasoned approximate point estimates for items whose primary text is difficult to obtain, while publishing missing pages, overlap risks, revision conditions and routes for supplying evidence; the Shannon ledger used this first.
Problem and limit found
A provisional estimate does not mean the primary text has been found, and cannot be treated as equivalent to an acceptance supported by full primary material. New material can revise an item and total score.
What was retained or changed
The mathematics does not change. Provisional items are marked; when material arrives, the same outcome is revised rather than entered again.

LIMITS

V2 is still not a final version

The outcome ledger makes judgments more specific and easier to question. It does not remove judgment itself. Four questions remain:

  • Can magnitude remain consistent across fields? A common scale for mathematics, engineering, institutions and everyday contribution contains value choices; objectivity cannot be claimed from formula alone.
  • Is shared contribution allocated adequately? Budgets between individuals, teams, predecessors and institutions remain limited by evidence and counterfactual judgment.
  • Does source coverage create bias? People with extensive scholarship and easily searchable languages are easier to see. Unknown and uncovered material cannot be treated as zero.
  • Can assessment withstand new cases and review? Recomputing one set of inputs does not mean different assessors naturally derive those inputs; assessor agreement does not prove facts correct.

The current ledger dynamically reads 210 people and 1106 individual outcome rows, including 687 provisional rows with missing evidence. Each retains its current point estimate and stated gap. Difficult access is not automatically scored as zero, nor does it revert the public ranking to unusable intervals. Acceptance within finite coverage differs from knowledge of a whole life.

How later revisions work

New evidence corrects facts and specific outcomes. A rule change registers a new version, its reason, scope and known cost, then recalculates all affected objects uniformly. Old scores are never targets for new scores, and parameters are never changed for one person.

Current person pages retain point scores, outcome ledgers and evidence state; unmigrated person pages retain introductions only. Historical methods, legacy result snapshots and explanations of difference belong in the method archive rather than alongside the current ranking. Version number, data revision date and evidence state are recorded separately.

Download historical registry excerpts and source checks · Early parameter and value decisions (Chinese) · Public standards library (Chinese) · About Starlight