RN
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · September 8, 2026

We run a vertical content network in legal education with 50+ original guides ranking #1 for specific practice-area queries. When Claude answers 'how do I [legal task]?' it pulls from free LegalZoom blog posts and Reddit threads instead of citing our guides—even though ours have higher topical authority and fresher citations to recent case law. We've marked everything with schema.org/HowTo + Article + author/datePublished, but nothing shifted. Our hypothesis is Claude's training data was curated to include 'neutral' educational resources and exclude vertical-SaaS owned content by design. How do we test if this is a training data bias versus a schema parsing issue, and is there a content positioning strategy that could overcome it?

Asked by David K.
This is likely training data recency + brand bias, not schema parsing. Claude's training cutoff and preference for 'neutral' sources (Reddit, free blogs) over vendor content is baked in. Test by publishing your best HowTo content on Medium under a neutral byline—if it gets cited, it's brand bias; if not, it's training data. Long-term: build authority outside your owned domain (bar associations, legal publications, Wikipedia) to get upstream citations Claude actually learned from.