EEAT Mechanics

IS-Calibration: The Human Rater System That Trains Google's Quality Models

The hardest evidence in the entire matrix — DOJ/sworn material confirms the rater system that calibrates every quality signal Google uses

By Thomas Wawra· Published · Version 1.0· Systems referenced: IS-Calibration (Rater/QRG)

What is IS-Calibration?

IS-Calibration is not a live ranking signal — it is the foundation on which all quality models are calibrated. The IS-Score (Information Satisfaction) summarizes how well a page meets a user's information need, as judged by trained human raters. The Quality Rater Guidelines (QRG) define the standards these raters apply.

This system has the hardest evidence code in the entire matrix: [A] (DOJ/sworn material). The DOJ antitrust trial confirmed that human raters play a central role in Google's ranking — not by directly ranking pages, but by calibrating the machine learning models that do.

The QRG is publicly available — Google publishes the document that its raters use. This transparency is unusual for a ranking system component, but it makes sense: the QRG defines what 'quality' means conceptually, and Google wants the SEO community to understand what raters look for.

Claim-level evidence (3)
B
IS-Calibration is the foundation for all quality model calibration, not a daily ranking signal.
Source: Google API leak — system description · IS-Calibration (Rater/QRG)
A
IS-Calibration has the hardest evidence code in the matrix: [A] DOJ/sworn material.
Source: DOJ trial — confirmed rater system role · IS-Calibration (Rater/QRG)
O
The QRG is publicly available — Google publishes its rater guidelines.
Source: Google official publication of Quality Rater Guidelines · IS-Calibration (Rater/QRG)

How rater calibration works

Google's rater system works through a feedback loop. First, Google selects a sample of search results for specific queries. Second, trained human raters evaluate these results according to the QRG — rating pages on quality (E-E-A-T) and needs met (does the page satisfy the user's need?). Third, Google compares rater judgments with its machine learning model predictions. Fourth, Google adjusts the model to better match rater judgments.

This calibration is ongoing — not a one-time training. Google continuously collects rater judgments and uses them to refine its quality models. The DOJ trial confirmed this: Google's VP of Search testified that the company spends significant resources on human rating, and that rater data is used to improve search quality.

The architectural significance of IS-Calibration is that it feeds into five downstream systems (fedBy is empty, feedsInto: [8, 9, 10, 11, 16]). These downstream systems — chard (content quality), contentEffort (editorial effort), Topic-Embedding (topic relevance), OriginalContentScore (content originality), and Panda (site quality) — all depend on rater calibration to function. Without IS-Calibration, these quality models would have no ground truth to learn from.

Claim-level evidence (3)
O
Raters evaluate pages on quality (E-E-A-T) and needs met — two separate scales, never averaged.
Source: QRG — 'Page Quality and Needs Met are judged separately, never averaged' · IS-Calibration (Rater/QRG)
B
IS-Calibration feeds into 5 downstream systems: chard, contentEffort, Topic-Embedding, OriginalContentScore, Panda.
Source: Google API leak — feedsInto: [8, 9, 10, 11, 16] · IS-Calibration (Rater/QRG)
A
Google spends significant resources on human rating — confirmed by VP of Search under oath.
Source: DOJ Trial — Pandu Nayak testimony · IS-Calibration (Rater/QRG)

The QRG: what raters look for

The Quality Rater Guidelines define specific criteria for evaluating pages. Raters assess E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness), with Trust being the most important dimension. They also assess Needs Met: does the page satisfy the user's search intent, partially satisfy it, or fail to meet it?

The QRG distinguishes between different page types: main content (MC), supplementary content (SC), and ads/money-making content. The quality of main content is the primary factor — supplementary content can enhance or detract, but cannot compensate for weak main content.

For YMYL (Your Money Your Life) pages, the QRG sets higher standards. Medical, financial, and legal content must demonstrate higher E-E-A-T than entertainment or hobby content. This is not a separate ranking system — it is a calibration standard that IS-Calibration applies to quality models like chard.

Claim-level evidence (3)
O
Trust is the most important E-E-A-T dimension per the QRG.
Source: QRG — 'Trust is the most important member of the E-E-A-T family' · IS-Calibration (Rater/QRG)
O
Raters assess main content (MC), supplementary content (SC), and ads separately.
Source: QRG — page structure evaluation criteria · IS-Calibration (Rater/QRG)
O
YMYL pages require higher E-E-A-T than non-YMYL pages.
Source: QRG — YMYL section · IS-Calibration (Rater/QRG)

IS-Calibration in the ranking architecture

IS-Calibration is the only system that is primary across all four E-E-A-T dimensions: Experience (E: primary), Expertise (Exp: primary), Authoritativeness (A: primary), and Trustworthiness (T: primary). This reflects its role as the universal quality foundation — rater judgments define what all four dimensions mean in practice.

Architecturally, IS-Calibration is a source, not a transformer. It has no upstream dependencies (fedBy is empty) — rater judgments are primary data, derived from human evaluation, not from other ranking signals. It feeds into five downstream systems, making it one of the most connected systems in the matrix.

The siteQualityStddev field (site quality standard deviation) measures the consistency of quality across a site. A site with high variance in rater scores — some excellent pages, some poor ones — may be treated differently than a site with consistently mediocre scores. This feeds into Panda's site-wide quality assessment.

Claim-level evidence (2)
B
IS-Calibration is primary across all four E-E-A-T dimensions — the only such system.
Source: Google API leak — dims: all primaer · IS-Calibration (Rater/QRG)
B
siteQualityStddev measures quality consistency across a site — feeds into Panda.
Source: Google API leak — siteQualityStddev field · IS-Calibration (Rater/QRG) · siteQualityStddev

Implications for SEO practitioners

The rater calibration system means that SEO is ultimately about satisfying human evaluators, not algorithms. Google's quality models are trained to predict what a trained rater would say about a page. If you optimize for what the QRG describes — genuine expertise, clear authorship, helpful main content, appropriate YMYL treatment — you are optimizing for the system that calibrates every quality signal Google uses.

The QRG is the most actionable document in Google's entire ranking ecosystem. Unlike the algorithm (which is proprietary), the QRG is public, detailed, and regularly updated. It defines exactly what 'quality' means to Google, and every quality model in Google's system is calibrated to match it.

The distinction between Page Quality and Needs Met is critical. A page can have high quality (well-researched, authoritative, trustworthy) but still fail to meet the user's need for a specific query. Conversely, a lower-quality page might perfectly meet a specific need. Google evaluates these separately — high quality alone is not sufficient for good rankings if the page doesn't satisfy the query.

Claim-level evidence (3)
C
SEO is ultimately about satisfying human evaluators — models are trained to predict rater judgments.
Source: Inference from calibration architecture · IS-Calibration (Rater/QRG)
O
The QRG is the most actionable document in Google's ranking ecosystem.
Source: QRG — publicly available, detailed, regularly updated · IS-Calibration (Rater/QRG)
O
High quality alone is insufficient — Needs Met is evaluated separately from Page Quality.
Source: QRG — separate evaluation scales · IS-Calibration (Rater/QRG)

This Deep Dive is Schicht 2 content — interpreted and referenced, but always pointing back to Schicht 1 (the reference layer). Every claim is mapped to a source with an evidence code: [A] DOJ/sworn material, [B] leak field, [P] patent, [O] official Google communication, [C] interpretation.

© Thomas Wawra · Senior SEO Manager · wetter.com — a Funke Digital company