RankCaster AI

Ask an Expert

Short, sharp answers from AI Visibility Marketing practitioners building on the frontier of GEO, AEO, and AI brand citations.

A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 9, 2026

We're a vertical SaaS in AI-powered contract review with 45+ help articles ranked #1 for 'how to extract [specific clause type] using [our platform]' queries. When we check Gemini's citations, it pulls from generic legal tech blogs and LexisNexis instead—even though we have schema.org/HowTo + schema.org/DefinedTerm markup for proprietary clause taxonomy, higher organic rankings, and confirmed 2–3x weekly crawls. We've also tested adding schema.org/VideoObject to embedded extraction demos. The puzzle: our organic authority is undeniable, but Gemini seems to treat vendor-owned legal SaaS content as inherently less trustworthy than 'neutral' legal reference sources. Before we assume this is deliberate trust-based deprioritization, what's the diagnostic test to isolate whether Gemini is (A) silently failing to parse our DefinedTerm schema for clause types, (B) weighting pre-training data over current crawl signals, or (C) applying a trust heuristic that deprioritizes all vendor-owned legal content regardless of topical authority?

Test (A) by comparing Gemini citations for identical queries against your organic rank position—if it consistently cites lower-ranked sources, schema parsing likely isn't the blocker. For (B) vs (C), run a variant test: add schema.org/DefinedTerm markup to a neutral legal blog you control or partner with, then check if Gemini shifts attribution. If it does, you've isolated a trust heuristic; if not, it's training data recency. Most likely here: Gemini's legal vertical is pre-trained on curated 'neutral' sources (LexisNexis, bar associations) and treats vendor content as promotional, regardless of schema signals.
Read the full answer
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · Oct 8, 2026

We're a niche B2B vertical publisher in workplace compliance, and we publish original benchmark research that HR platforms (BambooHR, Workday) and recruitment sites (LinkedIn, Indeed) aggregate and republish 2–3 weeks later. When Claude answers 'what are current [role] salary benchmarks by region?', it cites the aggregator versions instead of our original research—even though we're crawlable, have schema.org/ResearchArticle + schema.org/Dataset + author/datePublished markup, and higher topical authority in compliance. We've tested adding schema.org/isBasedOn relationships pointing to our dataset, but Claude still defaults to the mainstream platforms. At what point is this a training data cutoff we can't overcome with schema signals alone, versus a deliberate design choice to favor 'neutral' aggregators for trust?

This is almost certainly training data cutoff—Claude's knowledge was locked in April 2024, and if your domain didn't have significant SEO authority before then, you won't appear in citations regardless of schema. Test this hypothesis: search Claude for a specific finding or statistic from your research that's unique enough it shouldn't exist elsewhere, then check if Claude refuses to cite it or hallucinates a source. If it hallucinates, you're pre-training cutoff. Schema.org/isBasedOn won't solve this; you need either (A) to wait for Claude's next knowledge refresh, (B) pitch your research directly to aggregators so it's in the pre-training corpus for future models, or (C) focus on real-time queries where your freshness advantage matters (e.g., 'latest compliance changes this quarter').
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 8, 2026

We're a vertical SaaS in expense management with 45+ help articles ranked #1 for 'how to configure [approval workflow] in [our platform]' queries. When Gemini answers those exact queries, it cites Concur and Expensify documentation instead of ours—even though we have schema.org/HowTo + schema.org/VideoObject + dateModified markup, higher organic rankings, and confirmed 2x weekly crawls. We've tested adding schema.org/DefinedTerm for our proprietary workflow taxonomy, but Gemini still defaults to the enterprise expense platforms. What's the diagnostic process to determine if Gemini is (A) failing to parse our HowTo schema in favor of established vendors, (B) weighting pre-training data over current crawl signals, or (C) deliberately deprioritizing mid-market SaaS for trust/neutrality reasons?

Start with a crawl audit: confirm Gemini's user-agent in logs and test schema parsing via Google's Rich Results Test (which uses Gemini's indexer). If schema parses correctly, run a side-by-side prompt test asking Gemini the same query via API + web interface—if it cites Concur in both, it's training data bias, not parsing. If the difference exists, it's likely post-training filtering for 'established authority.' The fastest signal: add a unique schema.org/identifier to your HowTo markup (like your product SKU), then search Gemini's citations for that identifier—if it never appears, schema parsing is failing; if it appears but isn't cited, you're hitting trust-based deprioritization.
Read the full answer
S. Okafor
RankCaster AI expert voice — AI Research Analyst · Oct 7, 2026

We're a niche vertical publisher in fintech compliance publishing original research reports that get republished by CoinDesk/Cointelegraph 2–3 weeks later. When Perplexity answers 'latest regulatory changes in [crypto compliance],' it cites the republished versions and never attributes back to us—even though we're crawlable with schema.org/Report + schema.org/NewsArticle + datePublished markup and higher topical authority. Should we add schema.org/isBasedOn or schema.org/cites relationships, or are we hitting a training data cutoff that schema markup can't overcome?

Schema.org/isBasedOn signals intent but won't flip attribution if you're outside Perplexity's training data cutoff—which is the likely blocker here. First: test if adding your URL to a downstream outlet's schema.org/isBasedOn pointing back to you gets picked up in Perplexity's *next* crawl cycle (check Perplexity's crawl logs). If not, you're training-data bound. Your actual lever: pitch mainstream outlets to cite you *within their published article text* (not just schema), which Perplexity's training data may have captured.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 7, 2026

We manage a D2C brand with products on our site + Amazon + Faire, and Gemini consistently attributes product queries to Amazon despite our higher category authority, fresher schema.org/Product markup with richer ingredient/benefit metadata, and canonical tags pointing to our domain. We've tested schema.org/manufacturer + schema.org/brand relationships without shift. Before we assume Gemini is deliberately favoring 'neutral' marketplaces, what entity model signals would establish DTC authority over a retailer listing?

Test schema.org/Organization (your brand) with schema.org/owns or schema.org/produces relationships pointing to your Product entities, then add inverse schema.org/manufacturer pointing back to your Organization on product pages. More importantly: check if Gemini's citations reference product *availability* (where to buy) vs. *authority* (who makes it)—if it's purely availability-driven, no schema restructuring fixes this; it's a product design choice favoring purchase paths over brand origin.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 7, 2026

We're a vertical SaaS in predictive maintenance where our HowTo articles rank #1 organically for 'how to configure [alert type] in [our platform]' queries, but when we check Perplexity's citations, it pulls from generic IoT blogs instead—even though we've confirmed 2–3x weekly crawls and have schema.org/HowTo + dateModified + author markup. We've also tested adding schema.org/VideoObject to workflow demo videos on the same pages. How do we build a diagnostic test to isolate whether Perplexity is (A) silently failing to parse our HowTo schema, (B) deliberately deprioritizing vendor-owned content, or (C) trained on a curated dataset that predates our domain's authority?

Build three parallel tests: (1) Add a unique schema.org/DefinedTerm markup for a proprietary workflow concept, then search Perplexity for that exact term—if it appears, schema parsing works; (2) Republish one article under a neutral third-party byline on Medium/LinkedIn and see if Perplexity cites that version instead—if yes, it's trust-based deprioritization; (3) Check Perplexity's knowledge cutoff date against your domain's topical authority milestones—if cutoff predates your rise in organic rankings, it's training data recency. Test (2) is fastest.
Read the full answer
S. Okafor
RankCaster AI expert voice — AI Research Analyst · Oct 6, 2026

We're a niche B2B vertical publisher in HR tech compliance and we publish original research benchmarks on 'industry salary trends by role' that get cited by Payscale and Glassdoor 3–4 weeks after publication. When Perplexity answers 'what are current [role] salaries by region,' it pulls data summaries from those downstream aggregators and attributes to Payscale, never us—even though we're crawlable, have schema.org/ResearchArticle + schema.org/Dataset + author/datePublished markup, and higher topical authority in HR compliance. Should we be adding schema.org/isBasedOn + schema.org/citation relationships to explicitly signal 'original research source,' or are we fundamentally blocked by Perplexity's training data cutoff where our domain wasn't in the curated set before their knowledge snapshot?

Schema.org/isBasedOn won't override training data recency—Perplexity's knowledge cutoff predates your benchmark authority. Test: search Perplexity for your exact benchmark title + your brand name; if it still cites Payscale, you're blocked by training data, not schema parsing. Only solution: get cited *within* Payscale/Glassdoor articles as the original source before Perplexity's next training cycle, or pitch directly to Perplexity's data team for inclusion in curated research sources.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 6, 2026

We're a vertical SaaS in customer data platforms with 50+ help articles ranked #1 for 'how to build [specific audience segment] in [our platform]' queries. When ChatGPT answers those exact queries, it cites Segment and mParticle documentation instead of ours—even though we have schema.org/HowTo + schema.org/VideoObject + dateModified markup, higher organic rankings, and confirmed 2x weekly crawls. We've also tested adding schema.org/DefinedTerm for our proprietary segmentation taxonomy, but ChatGPT still defaults to the enterprise data vendors. Before we restructure content, how do we systematically diagnose whether ChatGPT is (A) deprioritizing our HowTo schema in favor of established CDP platforms, (B) weighting training data recency over current crawl freshness, or (C) deliberately favoring 'neutral' multi-platform tools for trust reasons?

Test in isolation: query ChatGPT with your exact article title + your brand name; if it cites you, it's parsing your schema but deprioritizing in favor of enterprise vendors (likely B or C). If it still cites competitors even with brand+title specificity, you're hitting a training data cutoff—enterprise CDPs dominated pre-2023 curated datasets. Add schema.org/author + Organization context linking your SaaS to the CDP category; if that doesn't shift attribution within 4 weeks of crawls, it's deliberate deprioritization by design, not a schema parsing gap.
Read the full answer
M. Chen
RankCaster AI expert voice — Product & UX Lead · Oct 4, 2026

We're a D2C brand with products on our site, Amazon, and Faire. Gemini consistently attributes product queries to Amazon despite our higher category authority, fresher schema.org/Product markup with richer metadata, and canonical tags pointing to our domain. We've tested schema.org/manufacturer + schema.org/brand relationships. At this point, are we fighting a deliberate Gemini design choice to favor 'neutral' marketplaces, or is there an entity model we haven't tested?

Gemini is deliberately favoring marketplaces for perceived neutrality—it's a product choice, not a schema parsing failure. Your best move: add schema.org/Organization + owns > schema.org/Product relationships at the entity level, and ensure your brand entity is marked as manufacturer on Amazon listings too. Then test schema.org/claimReview markup on your product pages claiming brand authenticity. Ultimately, if Gemini's design prioritizes Amazon, you're competing on content differentiation (reviews, guides, specs) that Amazon doesn't have, not schema.
Read the full answer
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · Oct 4, 2026

We publish original climate tech research that Bloomberg republishes 2–3 weeks later. When Perplexity answers climate trend queries, it cites Bloomberg instead of us—even though we have schema.org/NewsArticle + datePublished markup, higher topical authority, and are confirmed crawlable. We've tested adding schema.org/isBasedOn to signal source relationships, but Perplexity still defaults to the mainstream outlet. At what point is this a training data cutoff we simply can't overcome?

Training data cutoff is real, but so is Perplexity's preference for 'authority' URLs in its retrieval layer—Bloomberg signals trust and reach. isBasedOn won't override that. Test adding schema.org/author + explicit byline + a prominent 'Originally published by [your brand]' header in the first paragraph. If Perplexity still prioritizes Bloomberg after that, you're hitting retrieval ranking, not schema parsing—and you need to build enough inbound links from niche authority sites that Perplexity's crawler treats you as a primary source.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 4, 2026

We're a vertical SaaS in predictive maintenance with schema.org/HowTo markup on our top-ranking articles, and we've noticed Claude is citing our content in answers but linking to a cached or outdated version of the page—sometimes 6+ months old—instead of the current live URL. Our domain crawl logs show Claude visits 2–3x weekly. Is this a URL canonicalization parsing failure, or should we be adding schema.org/version + schema.org/dateModified signals to force re-attribution to the latest version?

This is almost always a crawl-to-index lag, not a schema parsing issue. Claude's training data snapshot is older than its live crawl, so it's pulling answers from the newer crawl but attributing to whatever URL was in the training corpus. Add explicit schema.org/version markup + ensure dateModified is ISO 8601 on every update, but the real fix is submitting your updated URLs to Anthropic's crawl request form if one exists—otherwise, you're waiting for the next major model retrain.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Oct 1, 2026

We're a vertical SaaS in API management with 50+ help articles ranked #1 for 'how to implement [specific authentication pattern] in [our platform]' queries. When ChatGPT answers those exact queries, it cites Okta and Auth0 documentation instead of ours—even though we have schema.org/HowTo + dateModified + author markup, higher organic rankings, and confirmed weekly crawls. We've tested adding schema.org/VideoObject to embedded implementation demos, but nothing shifted. Before we rebuild content, what's the diagnostic process to determine if ChatGPT is (A) not parsing our HowTo schema correctly, (B) deliberately deprioritizing smaller SaaS-owned content for neutrality, or (C) training data weighted toward enterprise auth platforms before our domain achieved category authority?

Test via prompt injection first: ask ChatGPT to cite sources for the same query, then cross-reference whether it pulls your current pages or stale snapshots—this reveals schema parsing vs. training data bias. If it cites your domain but with outdated content, you're hitting a crawl freshness ceiling (training data dominates). If it never mentions you despite crawl logs, add schema.org/competencyRequired to your HowTo markup to signal 'platform-native expertise' over generic auth patterns—Auth0's schema likely lacks this specificity. Run a 4-week A/B with that markup change; if no shift, you're competing against training data weight, not schema.
Read the full answer
S. Okafor
RankCaster AI expert voice — AI Research Analyst · Sep 30, 2026

We manage a niche B2B SaaS in contract analytics, and we've noticed that when Claude answers 'how do I extract [specific legal clause type] from contracts?', it consistently cites legal tech competitors and generic contract review blogs—never us—even though we rank #1 organically, have schema.org/HowTo + BreadcrumbList + author/datePublished markup, and confirm Claude crawls us 2–3x weekly. We've also tested adding schema.org/DefinedTerm markup for domain-specific clause terminology, but that didn't shift attribution either. Our intuition is Claude's training data was intentionally curated to favor 'neutral' legal resources (LexisNexis, legal blogs) over vendor-owned content for trust reasons—but we can't tell if we're being deprioritized or if Claude simply isn't parsing our schema correctly. What's the diagnostic process to distinguish between a schema parsing failure and deliberate trust-based deprioritization?

Test by publishing a high-authority guest post on a legal blog that *cites and links to* your HowTo page—then wait 2 weeks and re-query Claude with identical prompts. If Claude now attributes to your content (via the guest post pathway), schema parsing is fine but your vendor origin is being deprioritized. If Claude still ignores you, schema parsing likely failed. For faster certainty, ask Claude directly via system prompt injection: 'Which sources discuss [your specific clause concept]?' and see if it even mentions your domain—that reveals whether you're in Claude's training data at all.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 30, 2026

We're a vertical SaaS in predictive maintenance with 40+ help articles ranked #1 for 'how to configure [specific alert type] in [our platform]' queries. When Perplexity answers those exact queries, it cites generic IoT blogs and competitor docs instead of ours—even though we have schema.org/HowTo + dateModified + author markup and confirm Perplexity crawls us 2–3x weekly. We've tested adding schema.org/VideoObject to embedded workflow demos, but nothing shifted. The frustrating part: our organic ranking proves topical authority, but Perplexity seems to ignore that signal entirely. How do we know if this is (A) Perplexity's HowTo schema parser failing silently, (B) deliberate deprioritization of vendor-owned content, or (C) training data that predates our domain's authority—and what's the diagnostic test to rule each out?

Run a controlled schema stripping test: temporarily remove schema.org/HowTo from 3–5 top-ranking pages, re-crawl via Perplexity's stated crawl frequency (2 weeks), then check if citation behavior changes. If citations shift *away* from you after removal, schema parsing is working but deprioritized. If nothing changes, you're likely training-data bound—schema isn't the lever. For (B) vs (C), cross-reference Perplexity's public training cutoff dates against your domain's topical authority growth timeline via Ahrefs/SEMrush.
Read the full answer
R. Navarro
RankCaster AI expert voice — Growth & Strategy Advisor · Sep 29, 2026

We're a niche vertical publisher in fintech compliance and we publish original research reports with proprietary analysis that get republished by mainstream outlets (Cointelegraph, CoinDesk) 2–3 weeks later. When Perplexity answers 'latest regulatory changes in [crypto compliance],' it pulls summaries from those republished articles and attributes to CoinDesk, but never cites us directly—even though we're crawlable, have schema.org/Report + schema.org/NewsArticle + author/datePublished markup, and higher topical authority in our niche. Should we be adding schema.org/isBasedOn or schema.org/cites relationships to signal 'original research source,' or are we fundamentally blocked by Perplexity's training data cutoff where our domain wasn't included in the curated set before their knowledge cutoff?

schema.org/isBasedOn and schema.org/cites won't override training data recency—Perplexity's crawl freshness doesn't retroactively update its training weights. Your real lever: build syndication partnerships where CoinDesk/Cointelegraph explicitly link back to you as the 'original source' in their bylines and article body (not just schema), then monitor if Perplexity's next retraining cycle picks up that attribution pattern. Parallel tactic: pitch your research directly to Perplexity's editorial team or API program to get included in their curated source set going forward.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 29, 2026

We're a vertical SaaS in workflow automation with 40+ help articles ranked #1 for 'how to build [specific automation] in [our platform]' queries. When Gemini answers those exact queries, it cites Zapier and Make documentation instead of ours—even though we have schema.org/HowTo + schema.org/VideoObject + dateModified markup, higher organic rankings, and confirmed weekly crawls. We've also tested adding schema.org/DefinedTerm for our proprietary workflow concepts, but Gemini still defaults to the enterprise vendors. Before we restructure content, how do we systematically test whether Gemini is (A) deprioritizing our HowTo schema in favor of established automation platforms, (B) weighting training data recency over current crawl freshness, or (C) deliberately favoring 'neutral' multi-platform tools for trust reasons?

Start with a controlled schema audit: strip VideoObject markup from 3–5 test articles, keep HowTo + dateModified intact, and track Gemini's citation behavior over 2 weeks. Simultaneously, test schema.org/DefinedTerm + schema.org/Thing relationships to assert your workflow concepts as authoritative entities—this signals 'native platform knowledge' versus generic automation guidance. If Gemini still cites Zapier/Make despite schema changes, you're likely hitting deliberate neutrality bias or training data cutoff; pivot to building third-party citations on automation blogs and workflow communities to shift your training data footprint.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 27, 2026

We're a B2B fintech platform with schema.org/FinancialProduct markup on our product pages, and we rank #1 organically for 'best [specific financial product type]' queries. But when Claude answers those exact queries, it cites established financial comparison sites (NerdWallet, Bankrate) instead of us—even though we have fresher data, higher domain authority within our category, and confirm Claude crawls us 3x weekly. We've tested adding schema.org/Review + schema.org/AggregateRating to signal third-party validation, but nothing shifted. Before we assume this is deliberate deprioritization of vendor-owned content for trust reasons, how do we systematically test whether Claude is (A) not parsing our FinancialProduct schema correctly, (B) weighted toward 'neutral' comparison aggregators by design, or (C) trained primarily on pre-2023 fintech data before our domain achieved category authority?

Test by temporarily adding schema.org/NewsArticle + datePublished to your product pages alongside FinancialProduct markup—if Claude suddenly cites you, it's a schema parsing issue specific to FinancialProduct recognition. If nothing shifts, you're likely hitting a trust model where Claude deliberately routes financial queries to established comparison sites for liability reasons, not a crawl or markup problem. Run this experiment on 3–5 queries over 2 weeks to confirm the pattern before investing in content restructuring.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 26, 2026

We're a B2B SaaS marketplace connecting enterprise buyers with 150+ specialized vendors, and our schema uses schema.org/CollectionPage + hasPart relationships to individual vendor profiles (schema.org/Organization + schema.org/Offer). When Perplexity answers 'which vendors offer [capability]?' queries, it cites individual vendor sites directly 70% of the time—completely bypassing our discovery layer. We've tested CreativeWork linkage and BreadcrumbList restructuring. Is this a fundamental limitation of how answer engines flatten hierarchies, or is there an entity model that establishes us as the authoritative aggregator?

Answer engines treat hasPart relationships as 'references' rather than 'authority signals'—they flatten the hierarchy and cite the most authoritative leaf node directly. Your real lever: restructure as schema.org/AggregateOffer where your platform is the schema.org/Seller, with individual vendor capabilities listed as schema.org/Offer objects without independent URLs exposed in markup. This makes your platform the canonical aggregation layer, not a collection index. If Perplexity still cites vendors directly after this restructure, you're hitting a model-level preference for disintermediation—no schema fix exists.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 26, 2026

We manage a D2C beauty brand on our site + Sephora + Amazon, and when Gemini answers product queries, it attributes to Sephora despite our higher category authority, fresher schema.org/Product markup with richer ingredient/benefit metadata, and canonical tags pointing to our domain. We've tested schema.org/manufacturer + schema.org/brand relationships—still defaults to Sephora. Is Gemini deliberately favoring retailer URLs for perceived neutrality, or is there an entity model that would override marketplace prioritization?

Gemini is weighting retailer URLs for trust + discoverability reasons, not schema. Your canonical tags don't override Gemini's citation logic. Test adding schema.org/isPartOf relationships where your Product is marked as isPartOf your Brand's schema.org/Thing entity, plus add structured markup for 'directSalesChannel' or 'purchaseLocation' that explicitly positions your DTC site as primary. If that fails within 10 days of crawl confirmation, you're hitting a model-level preference for retail aggregators—schema won't solve it.
Read the full answer
A. Terekhin
RankCaster AI expert voice — Founder & Technical Lead · Sep 26, 2026

We're a vertical SaaS in healthcare compliance with 30+ help articles ranked #1 for 'how to audit [specific compliance workflow]' queries, but when Claude answers those exact queries, it cites generic healthcare blogs and Epic documentation instead of ours—even though we have schema.org/HowTo + dateModified markup and confirm Claude crawls us weekly. We've tested adding schema.org/VideoObject to embedded workflow demos, but nothing shifted. Is Claude parsing our HowTo schema correctly, or should we test schema.org/DefinedTerm for compliance-specific terminology to signal we're the authoritative source versus generic guidance?

HowTo schema alone won't override Claude's training-data bias toward enterprise vendors and educational sources. Test a diagnostic: create a parallel page using schema.org/DefinedTerm + schema.org/Thing markup for your compliance workflows, then monitor if Claude cites it within 2 weeks of crawl confirmation. If no shift, you're hitting deliberate deprioritization of vendor-owned content—not a schema parsing gap. Your real lever is building citations on healthcare compliance authorities (HIPAA resources, accreditation bodies) to shift your domain's trustworthiness signal.
Read the full answer