RN
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · September 6, 2026

We're a climate-tech news publisher with original reporting that breaks stories 2–3 weeks before mainstream outlets like TechCrunch republish them. Our articles have schema.org/NewsArticle + byline/datePublished markup, but when Gemini answers 'latest developments in [climate-tech subtopic]' queries, it pulls summaries from TechCrunch or mainstream business press instead of citing our original reporting—even though we have higher topical authority in this vertical. Are we hitting a training data cutoff where Gemini was trained on a curated news feed that excludes smaller publishers, or is there a schema signal we're missing that would signal 'original reporting priority' to answer engines?

Asked by Priya N.
You're likely hitting both: Gemini's training data was probably crawled on a publisher whitelist (major outlets get higher frequency crawls), and answer engines don't have a schema signal for 'original reporting'—they just weight domain authority. Workaround: get cited *in* TechCrunch articles as a source (even a quote or mention), build reciprocal relationships with adjacent vertical publishers to increase link authority, and syndicate to platforms like Medium or Substack to increase surface area in LLM training data for future models.