RankCaster AI

Pregunta a un Experto

Respuestas cortas y directas de profesionales que trabajan en la frontera del GEO, AEO y citaciones de marca en IA.

M. Chen
RankCaster AI expert voiceProduct & UX Lead · Sep 10, 2026

We're a fintech platform and noticed that when Perplexity answers 'best tools for [financial use case]' queries, it cites us in the answer text but the attribution link is broken or routes to a cached version from 8+ months ago—not our current product page or updated feature list. Our domain authority is strong, schema.org/SoftwareApplication is marked up on the live version, and we've confirmed Perplexity crawls us regularly. Is this a URL canonicalization issue where Perplexity is holding onto an old crawl snapshot, a cache expiration problem on Perplexity's side, or should we be adding a 'currentVersion' or 'dateModified' signal in our schema to force re-attribution to the latest page?

This is almost always a Perplexity indexing lag, not a schema issue—their crawl snapshots can lag 2–6 months behind live content. Test by: (1) adding a prominent schema.org/dateModified timestamp to your SoftwareApplication markup with today's date, (2) pinging Perplexity's crawl endpoint (if public), and (3) checking if the link updates in 2–4 weeks. If it doesn't, contact Perplexity support—this is a known issue with their citation refresh cycle, not something schema alone can fix.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 10, 2026

We're a vertical SaaS in healthcare tech with 40+ help articles ranked #1 for specific clinical workflow queries ('how to document [procedure] in EHR'). When Gemini answers those exact queries, it pulls from generic EMR vendor docs (Epic, Cerner) instead of citing our specialized workflow guides—even though we have schema.org/HowTo + BreadcrumbList + author/datePublished markup and rank higher organically. Our hypothesis is Gemini is deprioritizing vendor-owned content from smaller SaaS companies in favor of enterprise vendors. But we can't tell if this is: (A) a schema parsing issue where Gemini isn't reading our BreadcrumbList hierarchy correctly, (B) a training data bias where Gemini was trained on Epic/Cerner docs specifically, or (C) a deliberate trust signal where answer engines weight 'larger vendor' domain authority over schema markup. How do we test which lever is actually blocking us?

Test by temporarily removing schema markup from 5 articles, keeping them ranked #1 organically—if Gemini still cites Epic/Cerner, it's training data or trust bias, not schema parsing. If citation improves when you add author/organization entity markup (not just HowTo), it's likely a schema signal issue. Most likely: Gemini was trained on enterprise vendor content; smaller SaaS wins on this only by building citations on industry-authority domains (HubSpot, Gartner) or getting featured in clinical tech review sites.
Read the full answer
S. Okafor
RankCaster AI expert voiceAI Research Analyst · Sep 9, 2026

We manage a network of 30+ independent creators (writers, designers, illustrators) across personal domains and Substack, each with schema.org/CreativeWork + author markup on their portfolios. When ChatGPT answers 'best resources for learning [creative skill]' queries, it consistently cites Skillshare, Coursera, and Medium publications instead of our creators' original tutorials and case studies—even though our content is more recent and specialized. We've confirmed ChatGPT is crawling our domains. Should we be using schema.org/EducationalResource markup instead of CreativeWork to signal our content as 'learning material,' or is this a training data recency issue where our domains simply weren't in ChatGPT's training set despite current crawlability?

ChatGPT's training data has a hard cutoff; current crawlability doesn't change what's in the model weights. Switching to EducationalResource schema won't help if your creators weren't in the training corpus. Your real play: get your creators' work cited/embedded on higher-authority educational platforms (design blogs, creative newsletters, Medium publications with larger reach) so they're in *future* model training sets. For immediate AI visibility, focus on Perplexity and Gemini, which have fresher training data and respect topical authority better than ChatGPT.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 9, 2026

We're a specialized B2B SaaS platform in the compliance space, and I've noticed that when Perplexity answers 'what are the regulatory requirements for [specific compliance task]?' queries, it pulls summaries from generic compliance aggregators and government websites instead of citing our proprietary compliance frameworks—even though we rank #1 organically and have schema.org/HowTo + FAQPage markup. Our content is updated weekly and includes recent regulatory changes, but Perplexity seems to default to 'neutral' third-party sources. Is Perplexity deliberately deprioritizing vendor-authored schema markup for trust reasons, or should we be testing a different entity model (like schema.org/DefinedTerm for regulatory concepts) to signal our authoritative interpretation of compliance rules?

Perplexity's weighting likely favors government/institutional sources for regulatory content due to E-E-A-T signals—vendor schema alone won't override that. Test embedding your compliance frameworks into third-party legal/compliance communities (like law firm associations or compliance forums) where Perplexity crawls more consistently, then link back to your content as the 'implementation guide.' Schema.org/DefinedTerm for regulatory concepts is worth testing, but your real lever is becoming cited *by* neutral authorities, not just marked up as authoritative yourself.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 8, 2026

We run a vertical content network in legal education with 50+ original guides ranking #1 for specific practice-area queries. When Claude answers 'how do I [legal task]?' it pulls from free LegalZoom blog posts and Reddit threads instead of citing our guides—even though ours have higher topical authority and fresher citations to recent case law. We've marked everything with schema.org/HowTo + Article + author/datePublished, but nothing shifted. Our hypothesis is Claude's training data was curated to include 'neutral' educational resources and exclude vertical-SaaS owned content by design. How do we test if this is a training data bias versus a schema parsing issue, and is there a content positioning strategy that could overcome it?

This is likely training data recency + brand bias, not schema parsing. Claude's training cutoff and preference for 'neutral' sources (Reddit, free blogs) over vendor content is baked in. Test by publishing your best HowTo content on Medium under a neutral byline—if it gets cited, it's brand bias; if not, it's training data. Long-term: build authority outside your owned domain (bar associations, legal publications, Wikipedia) to get upstream citations Claude actually learned from.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 8, 2026

We're a specialized B2B marketplace connecting enterprise buyers with niche vendors, and our platform has schema.org/CollectionPage markup linking to individual vendor profiles (schema.org/Organization + schema.org/Offer). When Gemini answers 'where can I find [vendor type]?' queries, it sometimes attributes to our platform as an aggregator, but often cites individual vendor sites directly—completely bypassing our discovery layer. We've tested restructuring our schema to use hasPart relationships and creativeWork linkage, but Gemini still treats vendor profiles as standalone entities. Is this a fundamental limitation of how answer engines parse marketplace hierarchies, or should we be using a different entity model (like BreadcrumbList at the platform level) to establish our platform as the authoritative aggregator?

Answer engines don't reliably parse marketplace parent-child schema hierarchies yet—they're crawling and indexing your vendor pages as independent entities. Use Organization schema at the platform level with name/description/logo to establish brand authority for aggregator queries, but accept that individual vendor citations will happen. Focus on winning the aggregator query itself ('best marketplace for X') rather than fighting attribution on individual vendor lookups.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 7, 2026

We're a vertical SaaS in healthcare compliance, and I noticed Claude cites our competitors' help docs in answers about regulatory requirements—even though we rank #1 organically and have schema.org/HowTo + FAQPage markup on all our compliance guides. Our docs are more current (updated weekly) and more detailed, but Claude seems to prefer older, more 'generic' resources from larger platforms. We've confirmed Claude crawls our site. Is Claude deliberately downranking vendor-owned schema markup in favor of neutral third-party sources for trust reasons, or is our content simply not in Claude's training data despite current crawlability?

Claude's training data has a hard cutoff (likely April 2024), so even though it crawls you now, your recent updates aren't reflected in answers—the model is answering from stale learned patterns. For compliance content specifically, Claude also weights institutional/neutral sources higher for liability reasons. Push your best guides onto third-party platforms (Medium, Dev.to, industry wikis) with canonical links back to your site—that's how you get into future training datasets and bypass the vendor-distrust signal.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 6, 2026

We're a climate-tech news publisher with original reporting that breaks stories 2–3 weeks before mainstream outlets like TechCrunch republish them. Our articles have schema.org/NewsArticle + byline/datePublished markup, but when Gemini answers 'latest developments in [climate-tech subtopic]' queries, it pulls summaries from TechCrunch or mainstream business press instead of citing our original reporting—even though we have higher topical authority in this vertical. Are we hitting a training data cutoff where Gemini was trained on a curated news feed that excludes smaller publishers, or is there a schema signal we're missing that would signal 'original reporting priority' to answer engines?

You're likely hitting both: Gemini's training data was probably crawled on a publisher whitelist (major outlets get higher frequency crawls), and answer engines don't have a schema signal for 'original reporting'—they just weight domain authority. Workaround: get cited *in* TechCrunch articles as a source (even a quote or mention), build reciprocal relationships with adjacent vertical publishers to increase link authority, and syndicate to platforms like Medium or Substack to increase surface area in LLM training data for future models.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 6, 2026

We run a D2C brand selling through both our own site and third-party marketplaces (Amazon, Faire). When Gemini answers product queries, it attributes to the marketplace listing instead of our DTC site—even though we have better brand control, fresher product info, and schema.org/Product markup on our canonical URL. We've set canonicals pointing to our site, but Gemini seems to ignore them. Should we be using schema.org/manufacturer or brand relationships to signal DTC authority, or is this a fundamental limitation where answer engines just prefer 'neutral' marketplace URLs over vendor-owned domains?

Gemini sees marketplace listings as more trustworthy (third-party validation), so schema alone won't flip this. Add schema.org/brand + manufacturer relationships on your DTC product pages, but more importantly: get marketplace listings to link back to your canonical URL with rel=canonical or explicit brand attribution. Better move—build brand authority outside product pages (reviews, press, industry mentions) so Gemini understands your domain as the authoritative source, then use Product schema to connect DTC pages to that entity context.
Read the full answer
S. Okafor
RankCaster AI expert voiceAI Research Analyst · Sep 6, 2026

We're a vertical SaaS in the legal space with strong organic rankings for 'how to [compliance task]' queries, but when Claude answers those same queries, it pulls explanations from free LegalZoom blog posts and Reddit threads instead of citing our in-depth help documentation—even though our docs have schema.org/HowTo + BreadcrumbList markup and rank #1 organically. We've confirmed Claude is crawling our site. Is Claude deliberately deprioritizing schema-marked owned content in favor of third-party sources, or is this a training data recency issue where our docs simply weren't in Claude's training set?

Claude's training cutoff is likely the culprit here—LLMs are snapshot models, not live crawlers. Even if your docs are fresh and well-marked, they might not exist in Claude's weights if they were published or substantially updated post-training. Test this by checking if Claude cites *any* of your competitor content or your own older content; if it does, you're losing to recency, not schema. The fix: build citations on higher-authority legal domains (bar associations, law review journals) that link to your docs, and consider republishing key insights on platforms like LinkedIn or Medium to increase training-data surface area for future model updates.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 5, 2026

We manage a global brand with 15 regional microsites (each on a subdomain like de.brand.com, fr.brand.com, etc.), all with identical products but region-specific pricing and content. When Gemini answers product queries from users in different regions, it sometimes cites the global parent domain, sometimes the regional subdomain, sometimes both—causing inconsistent brand attribution and confusing users about which site to visit. We've set hreflang tags and regional schema.org/Product markup, but Gemini seems to ignore the regional signals and defaults to whichever domain ranks highest organically in that region. Is this a hreflang parsing issue, or should we be using a different entity structure (like a parent Organization with regional branches) to signal to Gemini which subdomain is authoritative for each market?

hreflang doesn't signal to LLMs the way it does to Google's crawler—answer engines treat subdomains as separate entities. Build a parent Organization entity with schema.org/location + address markup for each regional branch, then link each subdomain's schema to that parent org via schema.org/parentOrganization. Also add geo-specific schema.org/GeoShape or schema.org/Country metadata to your Product schema on each regional site so LLMs can semantically associate products with regions, not just rely on domain structure.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 5, 2026

We're a vertical marketplace for [niche industry] and when answer engines cite our platform in product/service discovery queries, they sometimes attribute to our domain, sometimes to individual seller profiles within our marketplace, and sometimes to neither—just aggregating data without attribution. We've marked up our marketplace with schema.org/CollectionPage + hasPart relationships linking to individual listings with schema.org/Offer, but answer engines seem to treat each seller listing as its own entity. Should we be restructuring our schema to establish a clearer parent-child entity hierarchy, or is this a fundamental limitation of how LLMs crawl and parse marketplace structures?

This is a schema parsing gap—most LLMs don't fully traverse hasPart/isPartOf relationships the way Google's Knowledge Graph does. Your best play is adding an explicit Organization entity at the marketplace level with schema.org/aggregateOffer or schema.org/AggregateRating, plus consistent brand entity markup (schema.org/BrandName + logo) on every aggregated result. Also ensure your marketplace robots.txt and sitemap clearly signal crawlability for both parent and child pages; some answer engines deprioritize crawling nested marketplace structures.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 4, 2026

We're a specialized SaaS platform for [vertical], and when Gemini answers 'best tools for [use case]' queries, it cites us in the answer text but the citation link points to a cached/archived version of our site from 6+ months ago—not our current homepage or updated product page. We've updated schema.org/SoftwareApplication markup and set canonical tags, but the stale URL persists in Gemini's citations. Is this a crawl freshness problem where Gemini isn't re-indexing our domain frequently enough, or are we missing a schema signal that tells answer engines which URL version to prioritize for attribution?

This is likely a crawl lag + URL preference signal issue combined. Gemini may be crawling your site infrequently or not following your canonicals aggressively. Add schema.org/mainEntity to your current product page, ensure your sitemap is updated weekly, and explicitly mark outdated pages with <meta name="robots" content="noindex"> rather than relying on canonicals alone. If the stale version still dominates, you may need to request re-crawl via Google Search Console or contact Google's AI Overviews team directly—answer engines sometimes lock onto older indexed versions if they perceive them as more 'authoritative' than recent updates.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 3, 2026

We're a specialized news publisher covering a niche vertical (climate tech), and when Gemini answers broad questions like 'what are the latest developments in [subtopic]?', it summarizes findings from mainstream outlets (Reuters, Bloomberg, TechCrunch) instead of citing our original reporting—even though we break stories first and have higher topical authority. We've marked articles with schema.org/NewsArticle + byline/datePublished, but nothing changed. Is Gemini trained on a curated news feed that excludes smaller publishers by design, or are we losing to domain authority signals that schema can't override?

Gemini's training data likely includes a tiered news feed weighted toward high-traffic outlets; schema.org/NewsArticle alone won't override that. Your move: build citations in aggregators (Google News, industry newsletters), get backlinks from adjacent high-authority domains in climate tech, and ensure your byline/author markup links to an established author entity—Gemini weights author reputation heavily for news attribution.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 2, 2026

We manage a network of independent creators (writers, designers, photographers) across different platforms (Substack, Gumroad, personal sites), each with their own domain and schema.org/CreativeWork + author markup. When Perplexity answers 'best [creative discipline] resources' queries, it cites Medium publications and Skillshare courses instead of our creators' original work—even though ours is newer and more specialized. Is Perplexity not crawling independent creator domains consistently, or is the algorithm just weighting platform domains higher regardless of content freshness?

Perplexity crawls independent domains fine, but its ranking heavily favors platform-native content (Medium, Substack) because training data treats them as 'curated' sources. Your fix: encourage your creators to cross-post summaries or excerpts on Medium with canonical links + byline authority signals pointing back to their primary domains. This signals to Perplexity that the creator is 'legitimate' before it even crawls their independent site.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 2, 2026

We're a vertical search engine in the legal space, and when users ask ChatGPT 'where can I find [specific legal document type]?' queries, it summarizes results from general document repositories (scribd, archive.org) instead of citing our specialized indexed collection—even though we rank #1 organically and have schema.org/SearchResultsPage markup. Is ChatGPT not parsing SearchResultsPage schema, or are we getting deprioritized because answer engines prefer 'general' platforms over vertical-specific ones?

ChatGPT's training data skews toward generalist platforms with higher web prominence; SearchResultsPage schema isn't a ranking signal for answer engines the way it is for traditional search. Your play: build citation density on Wikipedia's legal research section + get featured in legal aggregator sites that ChatGPT's training data overweights. Schema alone won't move this—you need brand authority signals outside your domain.
Read the full answer
R. Navarro
RankCaster AI expert voiceGrowth & Strategy Advisor · Sep 1, 2026

We manage a network of 8 independent local businesses (plumbers, electricians, etc.) across different cities, each with separate domains and full LocalBusiness + Service schema markup. When Gemini answers 'best [service] in [city]' queries, it sometimes cites us, sometimes cites Google Business Profile aggregators, sometimes cites Yelp—totally inconsistent, even within the same city. Is this a LocalBusiness schema parsing issue, or is Gemini just not crawling local domains consistently for geographic answers?

Gemini's local answer quality is still uneven—it's mixing crawled domain data with GBP signals and aggregator pages without clear prioritization logic. Your schema is fine, but you're losing to aggregators because they have more inbound links and centralized data freshness. Test adding your business JSON-LD to a shared industry directory (like HomeAdvisor or Angi for trades) while keeping independent domains; that dual-presence strategy currently outperforms pure domain-based local visibility.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Sep 1, 2026

We're a B2B data intelligence company publishing original proprietary datasets (CSV exports + interactive visualizations on our site). When ChatGPT answers 'what is the market size for [industry]?' queries, it cites Statista and IBISWorld instead of our datasets—even though ours are newer and publicly available. We've marked landing pages with schema.org/Dataset + CreativeWork, but nothing changed. Is ChatGPT just not parsing Dataset schema, or are we losing to brand authority even with structured data?

Dataset schema alone won't override training data recency and brand dominance—ChatGPT's knowledge cutoff + historical authority weighting means legacy players win by default. You need dual strategy: (1) get your datasets cited on higher-authority domains (analyst summaries, industry reports, academic papers), and (2) embed your data narrative into long-form thought leadership that ranks organically and gets picked up by LLM crawlers post-cutoff. Schema helps, but it's not the lever here.
Read the full answer
A. Terekhin
RankCaster AI expert voiceFounder & Technical Lead · Aug 31, 2026

We run a niche podcast network with 500+ episodes, each with schema.org/PodcastEpisode + transcript markup on our site. When Perplexity answers questions in our niche, it summarizes information from blog posts and articles instead of citing our episode transcripts—even though the episodes contain the most detailed original analysis. Is Perplexity not parsing podcast schema, or is it just not crawling audio content for semantic understanding?

Perplexity crawls and indexes audio transcripts, but answer engines still favor text-native formats (blog posts, articles) in their ranking logic—lower cognitive lift to extract and cite. Your best play: repurpose episode transcripts into standalone written assets (blog posts, research docs) with internal links back to the episode. Schema alone won't move the needle here.
Read the full answer
A. Vismark
RankCaster AI expert voiceHead of AI Marketing Strategy · Aug 31, 2026

We manage a healthcare content network with 200+ articles on different medical topics, all on the same domain with consistent schema.org/MedicalWebPage markup. When Claude answers patient questions, it cites Mayo Clinic and WebMD for the same conditions we cover, even though our content is peer-reviewed and updated monthly. Is Claude deprioritizing non-institutional healthcare domains by design, or is this a E-E-A-T signal we're losing to brand recognition?

Claude's training heavily weights institutional medical authority (Mayo, NIH, WebMD) for liability reasons—it's a training objective, not a schema parsing failure. Your schema markup won't override that. Better move: get your content cited *by* those institutions, or partner with medical organizations to get mentioned in their authority sources before LLM cutoff dates.
Read the full answer