EEAT Mechanics

What Good Content Is: contentEffort and Google's Soft Metrics

One leaked field, one rater definition, twelve official questions, and a checklist that marks which parts are evidence and which are our reading.

By Thomas Wawra· Published · Version 1.0· Systems referenced: contentEffort, NSR (+ Fallback inheritance), chard (+ YMYL/Hoax), IS-Calibration (Rater/QRG)

01One line in the leak

"What is good content?" has a surprisingly short machine answer in the leaked Content Warehouse documentation: a single field, contentEffort, described in one line as an LLM-based effort estimation for article pages. It sits in the module QualityNsrPQData, stored as a list of versioned values whose type is defined in Google's quality/nsr code, next to other page-level quality predictors such as chard, tofu, keto, rhubarb and vlq.

That one line is also everything the leak says. It names no inputs, no scale, no weight and no use in ranking, and it limits the field to article pages. A weather forecast, a calculator or a product page may never receive this value. Anything beyond the line is inference, and this article marks it as such.

Evidence per claim (4)
B
The leaked attribute description reads: "LLM-based effort estimation for article pages (see landspeeder/4311817)."[1]
B
contentEffort is typed as a list of QualityNsrVersionedFloatSignal, "A versioned float value" with a value and a unique version id, whose helper functions live in google3/quality/nsr/proto; several model versions of the score can therefore be stored side by side.[2]
B
The same module QualityNsrPQData holds a "URL-level chard prediction", a "URL-level tofu prediction", a "Keto score", a Rhubarb quality score and a "URL-level score of the VLQ model"; contentEffort is one page-level quality predictor among several.[3]
C
Because the description restricts the field to article pages and says nothing about inputs, scale, weight or ranking use, any claim about how contentEffort scores a given page, or that it scores non-article pages at all, goes beyond the artefact.[4]

02Google's own definition of effort

The word the field uses is the first word of Google's rater manual on content quality. The Search Quality Rater Guidelines of 11 September 2025 say that for most pages the quality of the main content can be determined by the amount of effort, originality, and talent or skill that went into it, plus accuracy for informational and YMYL pages. These four are the soft metrics: judgments a trained person makes, not numbers a tool reports.

The definition of effort is about people and purpose, not length or production method. Effort is the extent to which a human actively worked to create satisfying content, and it can go into building page functionality as much as into writing. What counts as enough depends on the page type, and the use of generative AI alone does not decide the level of effort.

Evidence per claim (4)
O
"For most pages, the quality of the MC can be determined by the amount of effort, originality, and talent or skill that went into the creation of the content. For informational pages and pages on YMYL topics, accuracy and consistency with well-established expert consensus is important."[5]
O
"Effort: Consider the extent to which a human being actively worked to create satisfying content." Effort may also go into page functionality, while running thousands of pages of free content through translation software "without any oversight, manual curation, etc., would not be considered to have effort."[6]
O
The guidelines tie the expected effort to the page: a short social media video needs less than a professionally produced documentary, "but both need sufficient effort to create satisfying content for their purpose"; raters are told to think about what effort looks like for the type of page.[6]
O
"the use of Generative AI tools alone does not determine the level of effort or Page Quality rating. Generative AI tools may be used for high quality and low quality content creation."[7]

03The soft metrics, spelled out

The rater manual gives each soft metric a high and a low end. High effort means well organised, edited and curated; low effort means helpful content mixed with filler, little discussion, or instructions buried under unrelated text. High originality means content unique to the site, own photos or footage, or a first-hand perspective; low originality means summaries of what others already said.

Two calibrations matter more than any single criterion. First, filler is named as its own failure: content that inflates a page without serving its purpose. Second, the bar is relative to the competition, because typical and average pages on a topic are Medium, not High. At the bottom, pages built almost entirely from copied, paraphrased or AI-generated material with no added value are rated Lowest, even when they credit their sources.

Evidence per claim (5)
O
High level of effort: "The MC is well-organized, edited, and curated to support the purpose." Low level of effort includes lack of curation or editing, where helpful content "is mixed with less helpful distracting or filler content", and lack of organization.[8]
O
High originality: MC unique or original to the website, original photos or video footage, or "a personal perspective based on first-hand life experience"; low originality: information "summarized from other sources with little added value" or summaries of others' product reviews.[9]
O
"Sometimes, MC includes "filler" - low-effort content that adds little value and doesn't directly support the purpose of the page. Filler can artificially inflate content, creating a page that appears rich but lacks content website visitors find valuable."[10]
O
Raters are told to "Have high standards!" and to calibrate against other pages on the same topic, because "typical" and "average" pages on a topic generally have Medium, not High, quality MC.[9]
O
The Lowest rating applies if all or almost all of the MC is copied, paraphrased, embedded, auto or AI generated, or reposted "with little to no effort, little to no originality, and little to no added value", even if the page credits another source.[7]

04Google's twelve questions for publishers

For publishers, Google translates the same ideas into self-assessment questions. The twelve content and quality questions ask about original information and analysis, completeness, insight beyond the obvious, added value over copied sources, honest titles, whether the page would be bookmarked or cited in print, value compared with other results, spelling, sloppy or hasty production, and mass production across many creators or sites.

Two further passages sharpen what effort is not and what it looks like. Google says outright that it has no preferred word count, and it asks publishers to show the "How": for reviews, how many products were tested, with what results and methods, backed by evidence such as photographs. On the other side of the line sits the spam policy on scaled content abuse, which applies no matter how the content is created.

Evidence per claim (4)
O
Google's content and quality questions include whether the content provides original information, reporting, research or analysis, whether it avoids simply copying or rewriting sources, and whether it is "produced well, or does it appear sloppy or hastily produced".[11]
O
"Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"[12]
O
Under "How", Google names as trust-building for product reviews the number of products tested, the test results and how the tests were conducted, "all accompanied by evidence of the work involved, such as photographs".[13]
O
Scaled content abuse is "when many pages are generated for the primary purpose of manipulating search rankings and not helping users", typically unoriginal content of little value "no matter how it's created"; examples include generating many pages with generative AI without adding value and stitching content from different pages without adding value.[14]

05Rater judgment and the machine estimate

Rater scores do not rank pages. Google states that no single rating can move a page, and that ratings serve to measure how well search works and to supply examples of helpful and unhelpful results. The rater's soft metrics are therefore not inputs; they are the yardstick against which systems are checked.

The leaked field and the rater manual share one word, and that is where the evidence ends. Nothing in the leak says that contentEffort is trained on rater judgments, and nothing in the guidelines mentions the field. Reading contentEffort as a machine approximation of the rater's effort criterion is plausible, and it is our interpretation.

Evidence per claim (2)
O
"No single rating can directly impact how a particular webpage, website, or result appears in Google Search"; instead, "ratings are used to measure how effectively search engines are working" and to provide "examples of helpful and unhelpful results for different searches".[15]
C
contentEffort and the rater criterion "Effort" share a name; neither artefact links them. Reading the field as an LLM approximation of the rater's effort judgment is an inference, not a documented training relationship.[16]

06The checklist

1. Purpose first: say what the page is for and put the most helpful part at the top. 2. Effort shows in the result: organised, edited, curated, with no filler ahead of the answer. 3. Something original: own data, own photos, own tests or a first-hand perspective; if you draw on sources, add substantial value. 4. Skill fits the task: the explanation works because the author can do the thing.

5. Accurate, and on YMYL topics consistent with expert consensus. 6. An honest title that neither exaggerates nor shocks. 7. Show the How: methods, tests, evidence, and a disclosure where automation did substantial work. 8. Calibrate against the results page: if your page is typical for the topic, raters call that Medium.

9. Do not write to a word count; length is not effort. 10. No scale without value: generating, stitching or paraphrasing pages in bulk is where the spam policy begins, whatever tool is used. Items 1 to 10 restate Google's documented criteria; that meeting them raises contentEffort is not documented anywhere.

Evidence per claim (1)
C
The checklist restates criteria from the rater guidelines and Google's publisher questions; none of them is documented as an input to contentEffort, so the list describes what Google says good content is, not how the field computes.[17]

07References

Quality Rater Guidelines (9)

  1. [5]OSearch Quality Rater Guidelines, 11 September 2025, section 3.2 "Quality of the Main Content"
  2. [6]OSearch Quality Rater Guidelines, 11 September 2025, section 3.2
  3. [7]OSearch Quality Rater Guidelines, 11 September 2025, section 4.6.6
  4. [8]OSearch Quality Rater Guidelines, 11 September 2025, section 7.1 "High Quality Main Content"
  5. [9]OSearch Quality Rater Guidelines, 11 September 2025, section 7.1
  6. [10]OSearch Quality Rater Guidelines, 11 September 2025, section 5.2.2 "Filler as a Poor User Experience"
  7. [15]OSearch Quality Rater Guidelines, 11 September 2025, section 0.1 "The Purpose of Search Quality Rating"
  8. [16]COwn comparison of QualityNsrPQData and the Search Quality Rater Guidelines
  9. [17]COwn synthesis of the Search Quality Rater Guidelines and "Creating helpful, reliable, people-first content"

This Deep Dive is Layer 2 content — interpreted and referenced, but always pointing back to Layer 1 (the reference layer). Every claim is mapped to a source with an evidence code: [A] DOJ/sworn material, [B] leak field, [P] patent, [O] official Google communication, [C] interpretation.

© Thomas Wawra · Senior SEO Manager · wetter.com — a Funke Digital company