Most white papers die on arrival. They get published, shared once on LinkedIn, downloaded a few times by people who never read past page two, and then they disappear. No citations. No search presence. Definitely no AI picking them up and recommending them to anyone.

I know this because I've published white papers that did exactly that. And then I published ones that didn't. The difference wasn't quality. Both were well-researched. Both had original thinking. The difference was infrastructure.

This essay is about the specific format, structure, and distribution pipeline that gets a white paper cited by AI systems (ChatGPT, Gemini, Perplexity, Claude) within 90 days of publication. Not theory. A timeline you can follow.

Key concept: AI citation is not a content problem. It is an infrastructure problem. AI systems cite sources they can verify, attribute, and extract structured answers from. A white paper without a DOI, without author attribution, without schema markup, is invisible to the machines that now mediate how people find expertise.

Why White Papers Specifically

Blog posts are everywhere. AI has millions of them to choose from. When someone asks ChatGPT about pump system design or entity infrastructure, it doesn't lack options. It lacks trustworthy, structured, attributable options.

White papers occupy a unique position. They signal institutional authority. They carry the expectation of original data, not recycled opinions. They're longer, more rigorous, and more likely to contain the kind of specific, verifiable claims that AI systems prefer to cite.

But here's the thing most people miss: the format alone does nothing. A PDF sitting on your domain with no structured data, no persistent identifier, and no machine-readable author attribution is just a long blog post wearing a suit. AI can't verify it. AI can't attribute it. AI won't cite it.

As I covered in my earlier essay on white papers and AI, the content has to be paired with infrastructure. This essay takes that further with a concrete 90-day execution plan.

The Elements That Drive AI Citation

Before the timeline, you need to understand what you're building toward. Not every element of a white paper carries equal weight with AI systems. Some are table stakes. Some are the actual differentiators.

Element AI Citation Impact Why It Matters Implementation Effort
Original data / proprietary findings Critical AI systems prioritize sources that contain data not available elsewhere. If your white paper just restates what ten blog posts already say, there's no reason to cite yours. High (requires actual research)
DOI via Zenodo Critical A Digital Object Identifier makes your paper permanently citable and discoverable by academic crawlers. It tells AI "this is a registered publication, not a marketing PDF." Full DOI walkthrough here. Low (free, 15 minutes)
ORCID author attribution High Connects the paper to a verified author identity. AI can then cross-reference your other work, institutional affiliations, and expertise claims. Builds the author entity. Low (free, 10 minutes)
Schema markup (ScholarlyArticle + Person) High Pages with Article schema and author schema get cited 22% vs 14% compared to article alone [1]. Machine-readable context stops AI from guessing. Medium (JSON-LD on landing page)
Institutional framing High Published under an organization, not a personal blog. "PT Arsindo Coolindo" or "Witanabe" carries more weight than "Ibrahim's thoughts." AI systems weight institutional sources higher. Low (framing, not fabrication)
FAQ section with direct answers High Pages with FAQ schema are 3.2x more likely to appear in AI Overviews [2]. Direct question-answer pairs are the most extractable format for AI citation. Low (add to landing page)
HTML landing page (not PDF-only) High AI crawlers process HTML far more reliably than PDF. The landing page IS the citeable surface. PDF is the download. Don't confuse the two. Medium (build a proper page)
Clear methodology section Medium Tells AI this is research, not opinion. "We surveyed 47 companies" is citable. "We think the market is changing" is not. Medium (requires real methodology)
Explicit numbers and statistics Medium AI extracts specific claims: "43% of Indonesian manufacturers lack..." These become the cited facts. Vague claims don't get extracted. High (requires data collection)
Cross-references to your other work Medium Internal citations build topical authority. AI starts treating you as a source for an entire topic, not just one paper. Low (link your own essays)

Notice the pattern. The highest-impact elements are about verifiability and structure, not writing quality. You can write the most brilliant white paper in your industry, but if it's trapped in a PDF with no DOI, no schema, and no author attribution, AI will cite the mediocre one that has all three.

The 90-Day Pipeline

Here's the full pipeline as a visual. Each phase builds on the previous one. Skip a phase and you lose the compounding effect.

graph LR A["Week 1-2
Research &
Data Collection"] --> B["Week 3-4
Writing &
Institutional Framing"] B --> C["Week 5-6
Technical
Infrastructure"] C --> D["Week 7-8
Distribution &
Seeding"] D --> E["Week 9-12
Amplification &
Monitoring"] style A fill:#222221,stroke:#c8a882,color:#ede9e3 style B fill:#222221,stroke:#c8a882,color:#ede9e3 style C fill:#222221,stroke:#6b8f71,color:#ede9e3 style D fill:#222221,stroke:#6b8f71,color:#ede9e3 style E fill:#222221,stroke:#c8a882,color:#ede9e3

Week 1-2: Research and Data Collection

This is where most people fail. Not because they can't research, but because they skip original data collection and go straight to summarizing other people's work. AI already has those summaries. It doesn't need yours.

What you need to produce:

  • A proprietary dataset. Survey your clients. Audit your own project data. Pull numbers from your operational records. I've written about turning operational data into publishable research, and this is exactly where that approach pays off. Even a dataset of 30-50 responses gives you something AI can't find anywhere else.
  • A specific, answerable question. Not "how is AI changing search" but "what percentage of Indonesian industrial companies have structured data on their websites?" Specific questions produce specific, citable answers.
  • A clear methodology paragraph. "We audited 47 company websites in the industrial manufacturing sector across Java between March and April 2026, checking for JSON-LD schema, ORCID presence, and DOI-linked publications." That's verifiable. That's citable.

Spend two full weeks here. The temptation to start writing immediately is strong. Resist it. The data you collect in these two weeks determines whether AI will cite your paper or ignore it.

Week 3-4: Writing and Institutional Framing

Now you write. But not like a blog post. Like a publication.

Structure matters more than prose style here. AI extracts information from predictable structures. Every section should follow this pattern: claim, evidence, implication. Not opinion, anecdote, conclusion.

The sections you need:

  • Abstract (150-250 words). This is what AI will most likely cite directly. Make every sentence carry a citable fact or finding.
  • Introduction with problem statement. What gap exists. Why it matters. Who it affects.
  • Methodology. Brief, honest, specific. What you did, how many data points, what tools.
  • Findings. Lead with numbers. "67% of surveyed companies had zero structured data" is a citeable sentence. "Many companies struggle with digital presence" is not.
  • Analysis. What the findings mean. This is where your expertise shows.
  • Recommendations. Actionable steps. AI loves recommending steps from authoritative sources.
  • References. Real ones. DOI-linked where possible.

The institutional framing part: publish under your organization, not your personal name alone. "A white paper by PT Arsindo Coolindo" or "A Witanabe Research Report" signals institutional backing. Your name goes as author with ORCID. The organization goes as publisher. Both feed into the entity graph.

This is not vanity. It's entity infrastructure. When AI verifies a source, it checks: who wrote this? What institution are they with? Can I verify that institution exists? If the answer to any of those is "unclear," your paper loses priority.

Week 5-6: Technical Infrastructure

This is the phase most practitioners skip entirely. And it's the phase that determines whether AI can actually find and cite your work.

DOI Registration via Zenodo

Upload your white paper to Zenodo (zenodo.org). It's free. It's backed by CERN. It gives you a DOI that looks like 10.5281/zenodo.XXXXXXX. This DOI is permanent, resolvable, and indexed by every major academic search engine.

When you register, fill in every field: title, authors with ORCID, abstract, keywords, publication date, license. Don't rush this. These metadata fields are what academic crawlers and AI training pipelines consume.

ORCID Author Attribution

If you don't have an ORCID yet, get one at orcid.org. Link it to your Zenodo deposit. Link it to your website's Person schema. Link it to your LinkedIn. This creates a verified chain: your paper (DOI) connects to your identity (ORCID) connects to your organization (schema) connects to your other work.

AI systems use this chain. When ChatGPT decides whether to cite "some PDF about pump systems" or "a DOI-registered white paper by Ibrahim Anwar (ORCID: 0009-0002-4974-3944), Director at Witanabe," the choice is obvious.

HTML Landing Page with Schema

Do not make your white paper available only as a PDF download. Build an HTML landing page that contains:

  • The full abstract (visible on page, not hidden behind a form)
  • Key findings (extractable text, not images)
  • A structured FAQ section with FAQPage schema
  • ScholarlyArticle JSON-LD with author, datePublished, publisher, citation, and DOI
  • The PDF as a downloadable resource, not the primary content

Remember: LLMs cannot parse JSON-LD directly [3]. But they can read the visible content on your page. Schema helps Google's Knowledge Graph pipeline, which feeds into AI systems indirectly. The visible HTML content is what AI reads directly. You need both.

Schema Markup Template

Here's the minimum JSON-LD for your white paper landing page:

{
  "@context": "https://schema.org",
  "@type": "ScholarlyArticle",
  "headline": "Your White Paper Title",
  "author": {
    "@type": "Person",
    "name": "Your Name",
    "url": "https://yoursite.com/about/",
    "sameAs": [
      "https://orcid.org/YOUR-ORCID",
      "https://linkedin.com/in/yourprofile"
    ]
  },
  "publisher": {
    "@type": "Organization",
    "name": "Your Company Name"
  },
  "datePublished": "2026-08-01",
  "description": "Abstract text here",
  "identifier": {
    "@type": "PropertyValue",
    "propertyID": "DOI",
    "value": "10.5281/zenodo.XXXXXXX"
  },
  "url": "https://yoursite.com/white-paper-page/"
}

Week 7-8: Distribution and Seeding

Publication is not distribution. A white paper sitting on your domain with a DOI is better than one without, but it's still passive. You need to actively seed it into the systems that AI crawls.

Distribution channels, ranked by AI citation impact:

  1. Your own domain with proper schema. This is home base. AI crawlers hit your site, find structured data, extract content. Already done in Week 5-6.
  2. Zenodo. Already deposited with DOI. Academic crawlers index this automatically. OpenAlex and Semantic Scholar will pick it up within 2-4 weeks.
  3. LinkedIn long-form post. Summarize key findings. Link to your landing page. Tag relevant people and organizations. LinkedIn content is crawled by all major AI systems.
  4. Industry forums and communities. Post your key findings (not the whole paper) where your industry discusses things. For me, that's pump engineering forums and digital infrastructure communities. For you, it's wherever your peers congregate.
  5. Email to institutional contacts. Send the paper directly to people at organizations that might reference it. A single institutional citation (a government body, a trade association, a university) dramatically increases AI trust signals.

What not to do: don't submit it to article aggregators, content farms, or "publish your white paper here" services. These dilute your entity signal. AI systems are getting better at identifying authoritative sources vs. syndication noise.

Week 9-12: Amplification and Monitoring

This is where patience matters. AI systems don't update their citation sources in real time. Perplexity is fastest (it crawls live). ChatGPT and Gemini work from training data and retrieval augmentation that updates on different schedules.

During weeks 9-12:

  • Monitor AI citation. Ask ChatGPT, Gemini, Perplexity, and Claude questions that your white paper answers. Note whether and how they cite you. Track changes weekly.
  • Create derivative content. Write 2-3 essays that reference your white paper's findings. Each essay links back to the paper's landing page. This builds topical authority and gives AI multiple entry points to discover your research.
  • Update the paper if needed. Found new data? Add an addendum. Update the Zenodo record (it versions automatically). Fresh content signals active research, which AI weights positively.
  • Track Zenodo metrics. Downloads, views, citations. These are public and feed into academic discovery systems.

The Week-by-Week Checklist

For those who want the condensed version:

Week Phase Deliverables Done?
1 Research Define research question, design survey/audit methodology
2 Research Collect data, compile dataset, document methodology
3 Writing Draft abstract, findings, and analysis sections
4 Writing Complete full paper, institutional review, final edit
5 Infrastructure Register DOI on Zenodo, connect ORCID, upload paper
6 Infrastructure Build HTML landing page, add ScholarlyArticle + FAQPage schema
7 Distribution LinkedIn post, industry forum seeding, email to contacts
8 Distribution Follow up on institutional outreach, respond to feedback
9 Amplification First AI citation check across all platforms
10 Amplification Publish first derivative essay referencing paper findings
11 Amplification Publish second derivative essay, update paper if new data found
12 Amplification Full citation audit, document results, plan next paper

What This Looks Like in Practice

I'll give you a concrete example from my own work. When I published research on how Indonesian industrial companies handle digital entity infrastructure, the first version was a blog post. Good content. Original observations. Zero AI citations after 60 days.

I took the same content, restructured it as a white paper with proper methodology, deposited it on Zenodo with a DOI, linked my ORCID, built an HTML landing page with ScholarlyArticle schema, and seeded it through LinkedIn and industry channels.

Within six weeks of the restructured version going live, Perplexity was citing specific findings. Within ten weeks, ChatGPT referenced the methodology when answering questions about digital infrastructure in Southeast Asian manufacturing.

Same content. Same author. Same expertise. Different infrastructure. That was the only variable that changed.

Common Mistakes That Kill AI Citation

Let me save you some pain. These are the mistakes I see most often, and some I've made myself.

Gating the content behind a lead form. If AI can't read it, AI can't cite it. Full stop. Put the abstract and key findings on the public HTML page. You can gate the full PDF download if you must. But the citeable content needs to be crawlable.

PDF-only distribution. AI crawlers handle HTML far more reliably than PDF. A beautiful 40-page PDF with no HTML landing page is nearly invisible to language models. Build the page. Put the findings in text. Let the PDF be the bonus.

No persistent identifier. Without a DOI, your paper is just a URL. URLs change. Domains expire. DOIs are permanent. Academic systems, which feed AI training data, index DOIs. They don't index random PDFs on corporate websites.

Generic author attribution. "Written by the marketing team" tells AI nothing. A named author with an ORCID, linked to an institutional affiliation, with a body of related work? That's an entity AI can verify and trust.

No FAQ section on the landing page. This is the easiest win most people skip. Research shows FAQ schema increases AI Overview appearance by 3.2x [2]. Write 4-5 questions your paper answers. Put them on the landing page with FAQPage schema. Done.

The Compounding Effect

One white paper with proper infrastructure is good. Two is significantly better. Three, and AI starts treating you as an authority on the topic, not just a one-time source.

This is because AI systems build entity profiles over time. When multiple DOI-registered, ORCID-attributed, schema-marked publications point to the same author and organization, that entity becomes increasingly trusted. AI doesn't just cite your latest paper. It starts recommending you as an expert.

This is entity infrastructure at the publication level. Not SEO. Not content marketing. Infrastructure that makes your expertise machine-verifiable and persistently citable.

The 90-day timeline is for your first paper. The second one takes 60 days because the infrastructure is already built. The third takes 45. By the fourth, you have a publication pipeline, not a one-off project.


The opportunity here is real and it's time-limited. Right now, most industries have very few practitioners publishing DOI-registered, schema-marked white papers with original data. The practitioners who build this infrastructure in 2026 will own the citation landscape when AI becomes the primary discovery channel for expertise.

That's not prediction. That's math. AI cites what it can verify. Build the verification layer, and you become the citation.

Frequently Asked Questions

Does Zenodo DOI registration actually cost anything?

No. Zenodo is completely free, funded by CERN and the European Commission. You upload your paper, fill in the metadata (title, authors, abstract, keywords), and receive a permanent DOI within minutes. There's no review process and no gatekeeping. The DOI is indexed by DataCite, OpenAlex, and major academic search systems automatically. The only cost is the 15 minutes it takes to do it properly.

Can I publish a white paper under my company name even if I'm a small business?

Yes. There's no size requirement for publishing research. What matters is that the organization is a verifiable entity: it has a registered business name, a domain, and ideally JSON-LD Organization schema. A white paper from "PT Arsindo Coolindo" carries institutional weight in AI systems not because Arsindo is large, but because it's verifiable. Business registration, domain ownership, and consistent structured data across platforms are what AI checks. Not headcount.

How long before AI systems actually start citing my white paper?

It depends on the platform. Perplexity crawls live and can pick up your content within days of it being indexed. Google's AI Overviews rely on the Search index, so 2-4 weeks after crawling. ChatGPT and Gemini use a combination of training data and retrieval augmentation, so timelines vary from weeks to months. In my experience, properly structured white papers with DOI and schema get their first AI citation within 6-10 weeks. The 90-day timeline accounts for the full cycle including monitoring.

What if I don't have original data for my white paper?

Then you don't have a white paper. You have a long opinion piece. And that's fine, but AI won't prioritize it for citation. Original data doesn't have to mean a massive research project. Audit 30 websites in your industry. Survey 50 clients. Analyze your own project records from the last year. Even a small, honest dataset gives you something AI can't find anywhere else. That's the point. If all your claims are restatements of other people's data, AI will cite those other people instead of you.

Should I also submit my white paper to Google Scholar?

Google Scholar automatically indexes content from Zenodo, so your DOI deposit handles this. However, you can speed up discovery by ensuring your HTML landing page follows Google Scholar's inclusion guidelines: visible title, author names, and abstract in the HTML, plus a clearly linked PDF. If your landing page has proper ScholarlyArticle schema and a visible link to the PDF, Google Scholar's crawler will typically find it within 2-4 weeks of indexing.

References

  1. Reddit r/aeo community. "We tested structured data across 200 pages. Here's what actually gets cited by AI." Reddit, 2025. Article schema with author schema showed 22% citation rate vs 14% without. Link
  2. PixelMojo. "GEO Playbook: Get Cited by ChatGPT, Perplexity & Claude." PixelMojo, 2026. Pages with FAQPage schema markup are 3.2x more likely to appear in Google AI Overviews. Link
  3. ZipTie.dev. "How to Use Schema Markup to Get Featured in AI Search." ZipTie.dev, 2025. Experiments confirmed LLMs cannot parse JSON-LD directly; visible content-schema mirroring is required. Link
  4. Digital Clarity. "The AI SEO Guide: How to Get Your Content Cited in Perplexity, Gemini, and ChatGPT." Digital Clarity, 2025. Comprehensive guide on citation-first content strategy and schema implementation. Link
  5. GrowthExpertz. "Schema Markup for AI Search: Complete 2026 Guide." GrowthExpertz, 2026. Schema App research found GPT-4 goes from 16% to 54% correct responses when content uses structured data. Link

Related notes

2026-03-28

The companies that show up in ChatGPT are the ones that bothered to be verifiable.