Reference · Ranking apparatus

The E-E-A-T mechanics, evidence by evidence

An evidence-disciplined reference of the 25 documented evaluation systems from the Content-Warehouse leak and DOJ material — their metrics and their contribution to Experience, Expertise, Authoritativeness, Trust.

Guiding principle

E-E-A-T is not a ranking signal and not a score in the code. It is the assessment frame of the Quality Rater Guidelines, against which part of the models are calibrated. This reference cleanly separates what is measured, what is flanked, and what is merely complied with.

~16,000
external Search Quality Raters worldwide

Google employs ~16,000 external raters who continuously evaluate search results according to the QRG — because quality, trust and relevance require human judgment that cannot be captured by algorithms alone. Raters represent real users across languages and cultures; their assessments calibrate the models that then rank billions of pages.

For LLMs
Raw version of this reference: /llms.txt ↗

All 25 systems, 38 leak fields and the best-practice mapping as plain text — to feed into Claude Code, Claude Skills, agents or RAG. Evidence codes [A]/[B]/[O]/[C] are preserved.

§A

The two axes

The rater logic separates two judgments — they are never combined into an average. Only beneath them lie the five function classes into which the systems sort.

Primary sources
Simplified QRG ↗
Accessible introduction to search quality evaluation — for non-raters.
Quality Rater Guidelines (full) ↗
The complete official handbook for Google raters — the standard behind E-E-A-T. Version 2025-09-11.
Ranking Systems Guide ↗
Official description of all named ranking systems — the foundation for §E and §F of this reference.
Helpful Content Guidance ↗
Google's official best practices for helpful, reliable content — the basis for §F of this reference.
Search Essentials ↗
The fundamental rules for appearing in Google Search.
Spam Policies ↗
What Google defines as spam and which practices lead to penalties.
Page Quality

How well the page fulfills its purpose — substance, creation effort, reliability. This is where the E-E-A-T frame lives. A property of the page itself, query-independent.

Needs Met

How well the page serves the concrete search intent. Query-dependent — not readable from the page alone, but only in relation to the query.

Never an average: Page Quality and Needs Met are judged separately.
Five roles in ranking
Every ranking system does one of five jobs — they are grouped by that job here.
BehaviorUser interaction — clicks and satisfaction as a re-ranking signal.n=5
Quality & PredictionContent quality and its model-based prediction.n=8
Penalty & SpamDeductions and manipulation defense.n=6
Index & InfrastructureProcessing, selection and serving — gates, not judgment.n=4
FreshnessRecency and its dating.n=2
§B

Where the systems pay in

Filled = primary contribution (measures/verifies directly, ●) · Outline = indirect contribution (flanks, ◐). Chip color = function class. Click a system to open its detail.

IS calibration (Rater / QRG) — the norm that defines all four dimensions · ●●●● · [A]
Not a runtime signal: it produces the target values the content models are trained against. It stands above the columns, not in one.
E
Experience
Self-experienced?
1 · ◐ 4
primary
indirect
E
Expertise
Expert-grounded?
2 · ◐ 9
primary
indirect
A
Authority
Recognized source?
4 · ◐ 8
primary
indirect
T
Trust
Reliable, honest?
10 · ◐ 9
primary
indirect
Foundation — Mustang / SuperRoot / Twiddler: scoring and serving infrastructure, evaluates nothing · ––––. Not optimizable, pays into no dimension.
§C

What each system derives its judgment from

Orthogonal to the function class. Provenance decides how you influence a dimension — and makes the influenceability gradient visible.

Influenceability ↓ ① Rater (steerable through substance) › ② Behavior (only via genuine satisfaction) › ③ Link (external attribution) › ④ Pattern (compliance only)
1Rater-based
Calibrated to human quality judgments (IS / QRG). The only provenance that addresses substance work directly.
chardcontentEffortOriginalContentScoretofu / ketoPandaQ*IS-Eichung
2Behavior-based
Calibrated to lived user demand (clicks, satisfaction). Not directly buildable — only earned through genuine satisfaction.
NavBoostGlueChromeRankBrainCrawl-Budget
3Link-based
From the link graph / external attribution. Barely a system of its own — flows primarily as a vector into NSR / siteAuthority. Penguin is the abuse side.
Anchor-Spam (Penguin)Link vector → NSR
4Pattern-based
Rule-, heuristic- and classifier-based (spam, dates, demotions). Not to be built up but to be complied with and avoided.
QualityBoost demot.SpamBrainhostAgeIP-PriorFreshnessTwiddlerDate triangul.
Language-based
Learned from language use itself (embeddings). Measures geometry / focus, not capability — a multiplier, not a substitute.
Topic-Embeddings
Composite & Infrastructure
Bundles several sources (NSR: links + behavior + focus + content) or judges nothing at all (Mustang, Tangram).
NSR (mixed)clutterScoreSegIndexerMustangTangram
The almost-empty Link band is the statement: in the documented material, links are barely a standalone evaluating system — they flow as a vector into the NSR composite. Dashed = not a standalone evaluating system (vector, composite part or infrastructure).
§D

Metric catalog

Every documented leak field individually: field name (monospace), definition, evidence badge and provenance. Searchable and filterable.

38/38 fields
Class
Provenance
Evidence
Behavior6 fields
goodClicks / badClicksB
Clicks of a query×page rated as satisfied or disappointed.
Behavior
lastLongestClicksB
Last, longest click of a session — the strongest satisfaction signal.
Behavior
GlueResponseB/C
Behavioral input on SERP features that Tangram assembles for display.
Behavior
chromeInTotal / uniqueChromeViewsB
Total Chrome views or unique viewers = site-wide attachment.
Behavior
unscaledIpPriorBadFractionB/C
Share of suspicious clicks per IP range (manipulation damping).
Pattern
voterTokenC
Identifies a click 'ballot', making repeat voting harder.
Pattern
Quality & Prediction14 fields
siteAuthorityB
Authority distilled from quality_nsr, applied in Q*.
Composite
lowQualityB
Low-quality flag, converted from quality_nsr.NsrData.
Composite
predictedDefaultNsrB
Historicized baseline quality score (VersionedFloatSignal → the trajectory counts).
Composite
nsrConfidenceB
Confidence in its own NSR judgment (deprecated).
Composite
chardScoreB
Content-quality value per document, calibrated to rater judgments.
Rater
chardVarianceB/C
Dispersion/uncertainty of the chardScore (high = an uncertain judgment).
Rater
contentEffortB
LLM-estimated creation effort and depth of a page.
Rater
OriginalContentScoreB/C
Originality/first-hand-material degree — the closest machine Experience proxy.
Rater
tofu / keto / RhubarbB/C
Codenames for refinement/delta signals at the subchunk level.
Rater
siteEmbeddingsB
Vector representation of the site's topic space.
Language
siteFocusScoreB
How sharply outlined the site's topic field is.
Language
siteRadiusB
How widely the site scatters in topic space (small = focused).
Language
IS-ScoreA
Information Satisfaction from rater judgments — the calibration norm.
Rater
siteQualityStddevB/C
Standard deviation of quality across a site (consistency).
Rater
Penalty & Spam10 fields
navDemotionB
Deduction for poor user guidance/navigation.
Pattern
anchorMismatchB
Anchor text does not match the link target.
Link
serpDemotionB
A deduction derived from SERP behavior.
Behavior
clutterScoreB
A measure of layout overload and distraction.
Pattern
scamnessB
Scam/fraud proximity of a page.
Pattern
unauthoritativeScoreB
Lack of authoritativeness.
Pattern
BabyPandaV2B/O
Site-wide thin/low-quality demotion (Panda heritage).
Rater
phraseAnchorSpamPenaltyB
Penalty for over-optimized anchor phrases.
Link
IsAnchorBayesSpamB
Bayes-classifier flag for anchor spam (yes/no).
Link
hostAgeB/C
Host age; young hosts receive less trust credit.
Pattern
Index & Infrastructure4 fields
scaledSelectionTierRankB/C
Rank for index-tier selection (0–32767 = 16-bit maximum).
Composite
crawl capacity × demandO
What the server can handle × how much Google wants to fetch the URLs.
Behavior
RankEmbedBERTB/O
BERT-based embedding for deeper query understanding.
Language
Mustang / Ascorer / SuperRootB/C
Scoring/serving infrastructure and compositor — evaluates nothing.
Infrastructure
Freshness4 fields
bylineDateB
A visibly stated date — cheaply faked.
Pattern
syntacticDateB
Extracted from URL/markup/timestamp — cheaply faked.
Pattern
semanticDateB
Inferred from the content — only fakeable through genuine updating.
Pattern
FreshnessTwiddler / RealTimeBoostB
Freshness boost (QDF) and short-lived real-time spikes.
Pattern
§E

System detail

All 25 systems with the matrix columns. Mechanics shows the technical internals; Plain text foregrounds the strategic consequence for wetter.com.

#SystemE · Ex · A · TEvidenceClass
§F

From Best Practice to Mechanism

Google's official Helpful Content best practices — mapped to the officially named ranking systems. The mapping itself is interpretive [C]; both ends are official [O].

Mapping = interpretive [C] · Official ends [O] · Leak fields [B]/[A] in optional deep-dive

Original information, own research or analysis; substantial, complete coverage of the topic that goes beyond the obvious — not mere summaries or rewrites.

Official backing
High
Rater involvementDirect
chard, OriginalContentScore and contentEffort are directly calibrated to rater judgments.
Official systems
Ensures first-hand material is surfaced; does not reward scraping or copying.
Evaluates people-first helpfulness site-wide.
Favored high-quality, original content.
Strongest cluster: substantial first-hand content feeds multiple official systems at once.
Technical deep-dive
Leak fields — not officially confirmed by Google.
OriginalContentScore[B]
contentEffort[B]
chardScore[B]

Content for people, not primarily for search engines; no mass or automated production without added value; no content that exists only to capture traffic.

Official backing
Medium
Rater involvementDirect
The Helpful content system and Panda are calibrated to rater quality judgments.
Official systems
Demotes search-engine-first content.
Addresses scaled content abuse among other issues.
The penalty side of the substance cluster: the same property viewed from the demotion angle.
Technical deep-dive
Leak fields — not officially confirmed by Google.
Panda / BabyPandaV2[B]
niedriger contentEffort (Risikomarker)[B]

The content satisfies what someone is actually searching for and leaves the feeling of being well served — clear search intent met.

Official backing
High
Rater involvementIndirect
Neural matching/RankBrain/BERT are algorithmic; raters confirm indirectly via the Needs Met rating.
Official systems
Understands meaning behind query and content.
Connects words with concepts.
Understands word combinations and intent.
Finds relevant individual passages.
The official side is query/meaning understanding; behavioral confirmation (NavBoost) is partly outside the text itself.
Technical deep-dive
Leak fields — not officially confirmed by Google.
NavBoost / Glue (goodClicks, lastLongestClicks)[B]Behavioral proxy for Needs Met — not produced by text alone, but earned through genuine satisfaction.

Cleanly produced: free of errors, not carelessly made, without intrusive advertising, usable on mobile.

Official backing
Medium
Rater involvementNone
Page Experience relies on technical signals — no direct rater involvement.
Official systems
Mobile-friendliness, HTTPS, Safe Browsing, no intrusive interstitials.
Directly addressable; provenance pattern.
Technical deep-dive
Leak fields — not officially confirmed by Google.
clutterScore[B]
scamness[B]

Recognizable expertise in the topic area; the site/author has a traceable background; others recommend or cite the source.

Official backing
Medium
Rater involvementDirect
Reliable information systems and IS-Score are built on rater judgments.
Official systems
Elevates authoritative pages, demotes low-quality ones, rewards quality journalism.
Reputation/endorsement through the link structure.
Partly only earnable over time and through external attribution, not by writing alone.
Technical deep-dive
Leak fields — not officially confirmed by Google.
NSR / siteAuthority[B]
Topic-Embeddings (siteFocusScore)[B]
IS-Score (Information Satisfaction)[A]

Clear who created the content and why; the responsible party/author is identifiable and trustworthy; for YMYL: accuracy and reliability as the core E-E-A-T component.

Official backing
Low
Rater involvementDirect
Strongest rater dependency: Google checks trust/accountability primarily through rater judgment — no algorithmic field replaces it.
Official systems
The only broad official system; demotes unreliable content, elevates authoritative.
KEY FINDING: most thinly backed by named official systems. Google checks accountability primarily through rater judgment — no checkable field. No marker strategy works here; genuine substance counts most.
Technical deep-dive
Leak fields — not officially confirmed by Google.
unauthoritativeScore[B]
Trust-Demotions (QualityBoost-Familie)[B]

For time-sensitive topics, content is current and maintained/updated as needed.

Official backing
Medium
Rater involvementNone
Freshness systems work algorithmically on date signals — no rater calibration.
Official systems
Surfaces fresher content where recency is expected (QDF).
Strategically central for wetter.com: time-critical forecasts and news.
Technical deep-dive
Leak fields — not officially confirmed by Google.
FreshnessTwiddler / RealTimeBoost[B]
semanticDate (teuer fälschbar)[B]

No manipulative practices: no keyword stuffing, no link schemes, no misleading techniques.

Official backing
High
Rater involvementNone
SpamBrain and Penguin are ML/pattern-based — no rater involvement.
Official systems
Detects spam patterns algorithmically.
Demotes spammy link building.
Prevents excessive credit for keyword-match domains.
Well backed; avoidance cluster (what NOT to do).
Technical deep-dive
Leak fields — not officially confirmed by Google.
phraseAnchorSpamPenalty / IsAnchorBayesSpam[B]
Context: two systems without a best-practice equivalent
Site diversity system [O]: max. ~2 results per domain in top results — a display rule, not a quality judgment. Consequence: consolidate to one strong URL per search intent.
Product reviews system [O]: rewards well-researched, original reviews (peripheral for wetter.com).
Almost every official best practice has mechanical counterparts — often several. But the density is uneven: where many systems underlie (substance, spam avoidance, intent), the guidance is hard-backed. Where few do (accountability/trust), Google relies on rater judgment and learned representations — no marker strategy works there, only genuine substance. Lead with the guidance; verify with the systems — never the reverse.
§G

System dependency chains

Select a system to see what feeds it (upstream) and what it feeds into (downstream). The rater connection shows whether Quality Rater judgments directly, indirectly, or conceptually inform this system.

DirectIndirectConceptual
Select system
← Select a system