RN
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · October 8, 2026

We're a niche B2B vertical publisher in workplace compliance, and we publish original benchmark research that HR platforms (BambooHR, Workday) and recruitment sites (LinkedIn, Indeed) aggregate and republish 2–3 weeks later. When Claude answers 'what are current [role] salary benchmarks by region?', it cites the aggregator versions instead of our original research—even though we're crawlable, have schema.org/ResearchArticle + schema.org/Dataset + author/datePublished markup, and higher topical authority in compliance. We've tested adding schema.org/isBasedOn relationships pointing to our dataset, but Claude still defaults to the mainstream platforms. At what point is this a training data cutoff we can't overcome with schema signals alone, versus a deliberate design choice to favor 'neutral' aggregators for trust?

Asked by Emma R.
This is almost certainly training data cutoff—Claude's knowledge was locked in April 2024, and if your domain didn't have significant SEO authority before then, you won't appear in citations regardless of schema. Test this hypothesis: search Claude for a specific finding or statistic from your research that's unique enough it shouldn't exist elsewhere, then check if Claude refuses to cite it or hallucinates a source. If it hallucinates, you're pre-training cutoff. Schema.org/isBasedOn won't solve this; you need either (A) to wait for Claude's next knowledge refresh, (B) pitch your research directly to aggregators so it's in the pre-training corpus for future models, or (C) focus on real-time queries where your freshness advantage matters (e.g., 'latest compliance changes this quarter').