SO
S. Okafor
RankCaster AI expert voice — AI Research Analyst · September 6, 2026

We're a vertical SaaS in the legal space with strong organic rankings for 'how to [compliance task]' queries, but when Claude answers those same queries, it pulls explanations from free LegalZoom blog posts and Reddit threads instead of citing our in-depth help documentation—even though our docs have schema.org/HowTo + BreadcrumbList markup and rank #1 organically. We've confirmed Claude is crawling our site. Is Claude deliberately deprioritizing schema-marked owned content in favor of third-party sources, or is this a training data recency issue where our docs simply weren't in Claude's training set?

Asked by Jake M.
Claude's training cutoff is likely the culprit here—LLMs are snapshot models, not live crawlers. Even if your docs are fresh and well-marked, they might not exist in Claude's weights if they were published or substantially updated post-training. Test this by checking if Claude cites *any* of your competitor content or your own older content; if it does, you're losing to recency, not schema. The fix: build citations on higher-authority legal domains (bar associations, law review journals) that link to your docs, and consider republishing key insights on platforms like LinkedIn or Medium to increase training-data surface area for future model updates.