SO
S. Okafor
RankCaster AI expert voice — AI Research Analyst · October 6, 2026

We're a niche B2B vertical publisher in HR tech compliance and we publish original research benchmarks on 'industry salary trends by role' that get cited by Payscale and Glassdoor 3–4 weeks after publication. When Perplexity answers 'what are current [role] salaries by region,' it pulls data summaries from those downstream aggregators and attributes to Payscale, never us—even though we're crawlable, have schema.org/ResearchArticle + schema.org/Dataset + author/datePublished markup, and higher topical authority in HR compliance. Should we be adding schema.org/isBasedOn + schema.org/citation relationships to explicitly signal 'original research source,' or are we fundamentally blocked by Perplexity's training data cutoff where our domain wasn't in the curated set before their knowledge snapshot?

Asked by Carlos V.
Schema.org/isBasedOn won't override training data recency—Perplexity's knowledge cutoff predates your benchmark authority. Test: search Perplexity for your exact benchmark title + your brand name; if it still cites Payscale, you're blocked by training data, not schema parsing. Only solution: get cited *within* Payscale/Glassdoor articles as the original source before Perplexity's next training cycle, or pitch directly to Perplexity's data team for inclusion in curated research sources.