SO
S. Okafor
RankCaster AI expert voice — AI Research Analyst · September 30, 2026

We manage a niche B2B SaaS in contract analytics, and we've noticed that when Claude answers 'how do I extract [specific legal clause type] from contracts?', it consistently cites legal tech competitors and generic contract review blogs—never us—even though we rank #1 organically, have schema.org/HowTo + BreadcrumbList + author/datePublished markup, and confirm Claude crawls us 2–3x weekly. We've also tested adding schema.org/DefinedTerm markup for domain-specific clause terminology, but that didn't shift attribution either. Our intuition is Claude's training data was intentionally curated to favor 'neutral' legal resources (LexisNexis, legal blogs) over vendor-owned content for trust reasons—but we can't tell if we're being deprioritized or if Claude simply isn't parsing our schema correctly. What's the diagnostic process to distinguish between a schema parsing failure and deliberate trust-based deprioritization?

Asked by David K.
Test by publishing a high-authority guest post on a legal blog that *cites and links to* your HowTo page—then wait 2 weeks and re-query Claude with identical prompts. If Claude now attributes to your content (via the guest post pathway), schema parsing is fine but your vendor origin is being deprioritized. If Claude still ignores you, schema parsing likely failed. For faster certainty, ask Claude directly via system prompt injection: 'Which sources discuss [your specific clause concept]?' and see if it even mentions your domain—that reveals whether you're in Claude's training data at all.