State-owned enterprises sit on a mountain of institutional authority. Government ownership. Decades of public record. Thousands of procurement documents. Regulatory filings stretching back to independence-era mandates. News coverage in the national press, every single quarter, for years.

AI systems love this kind of entity. Dense connections. High-trust sources. Multi-domain verification. An SOE is, structurally, the ideal entity for AI citation. The kind of entity that knowledge graphs treat as authoritative by default.

And yet. Ask ChatGPT about most SOEs and you will get a vague paragraph that sounds like it was scraped from a 2019 Wikipedia stub. Ask Gemini about a specific SOE project and you will get hallucinated details mixed with outdated facts. Ask Perplexity for the current leadership of a government-linked company and you will get the person who left three years ago.

This is not an AI problem. This is a structuring problem. SOEs have the authority. They just never bothered to make it machine-readable.

I have worked with government and SOE clients. I have seen the inside of their digital infrastructure. And I can tell you with confidence: the gap between what these organizations are and how AI systems represent them is the largest entity gap in any sector I have encountered.

Key concept: SOEs and government-linked companies possess the highest natural entity authority of any organization type. Government ownership, regulatory filings, institutional documentation, and decades of public record create exactly the kind of dense, verifiable entity profile that AI systems prefer. The problem is not authority. The problem is that none of it is structured for machine consumption. Leadership has no persistent identifiers. Projects have no DOIs. Organizational schema is absent. The result: AI systems default to stale, generic representations of institutions that deserve precision.

The SOE advantage that nobody is using

Let me be specific about what SOEs have that private companies would pay millions to manufacture.

Government registry presence. Every SOE is registered with the Ministry of State-Owned Enterprises (or equivalent). That means government databases, regulatory filings, annual reports submitted to parliament, and audit records. These are the exact document types that AI training pipelines weight heavily because they are institutional, verified, and persistent.

Cross-institutional connections. An SOE does not exist in isolation. It is connected to ministries, regulators, other SOEs, international bodies, bilateral agreements, and multilateral frameworks. In knowledge graph terms, this is extraordinary graph centrality. The more connections to other high-authority entities, the higher your own authority score [1].

Historical depth. Many SOEs have operational histories spanning decades. Some trace their lineage to colonial-era enterprises. That temporal depth creates a rich entity profile that AI systems use for confidence scoring. A company that has existed for 50 years, with 50 years of verifiable records, is treated very differently from a company founded last Tuesday.

Media coverage density. SOEs generate news coverage simply by existing. Quarterly earnings, leadership changes, government policy shifts, parliamentary inquiries. Every one of those news articles is a data point that AI systems ingest. Private companies spend fortunes on PR to generate this kind of coverage. SOEs get it automatically.

Procurement documentation. Government procurement creates paper trails that are, by regulation, public or semi-public. These documents live in government databases that are crawled by AI training pipelines. When an SOE procures services or completes a major project, that documentation becomes part of the permanent record.

I wrote about this dynamic in Why Institutional Clients Make Your Entity Unassailable. The authority transfer mechanism works in both directions. When your company works with an SOE, you inherit their authority. But the SOE also benefits from having a structured ecosystem of connected entities. The problem is that SOEs rarely structure the connections.

Where SOEs consistently fail

Here is what I keep seeing. The same pattern, across different countries, different sectors, different sizes of SOE. The failures are remarkably consistent.

Leadership has no persistent identifiers. The CEO of a major SOE has no ORCID. No Wikidata item. No structured data connecting them to the organization. When they publish a policy paper or give a keynote at an international conference, that contribution exists in a vacuum. AI systems cannot connect the person to the institution because there is no machine-readable link.

I wrote specifically about why ORCID matters beyond academia in ORCID Is Not Just for Academics. The persistent identifier infrastructure exists. It is free. SOE leadership simply does not use it.

Projects have no DOIs or structured documentation. An SOE completes a billion-dollar infrastructure project. Where is the structured record? A press release. Maybe a line item in an annual report PDF that no AI system can parse. No DOI. No structured dataset. No machine-readable project page with schema markup. The project might as well not exist for AI purposes.

Websites are institutional brochures, not entity hubs. Most SOE websites are built as digital brochures. Static pages with mission statements, organizational charts as JPEG images, and news sections that have not been updated since the last leadership change. There is no JSON-LD schema. No structured data for the organization, its leadership, its subsidiaries, its projects. The website exists for humans who already know to visit it. It does nothing for machines trying to understand the entity.

Brand mentions are unlinked. SOEs get mentioned constantly. In news articles, government reports, academic papers, industry publications. But as I discussed in Brand Mentions Without Links Are Wasted Authority, unstructured mentions without links or entity identifiers are leaking authority. AI systems see the mention but cannot confidently connect it to the canonical entity.

No sameAs declarations. An SOE might have a Wikipedia page, a Wikidata item, a Bloomberg profile, a government registry entry, and a LinkedIn page. None of these are connected with sameAs schema. Each exists as a separate, fragmented representation. AI systems have to guess that they are all the same entity. Sometimes they guess wrong.

SOE entity advantages vs implementation gaps

Entity Advantage What SOEs Have Current Implementation Gap Severity
Government registry Official registration, regulatory filings, parliamentary records Exists in PDF/paper form. Not machine-readable. No schema markup on registry pages. Critical
Cross-institutional connections Links to ministries, regulators, other SOEs, international bodies Mentioned in text. No structured sameAs or relatedLink declarations. Critical
Historical depth Decades of operational records, founding documents, milestones Timeline buried in About page. No schema foundingDate or historical event markup. High
Media coverage Quarterly press coverage, leadership news, policy announcements Coverage exists but brand mentions are unlinked. No structured press page. High
Leadership authority C-suite with government appointments, international conference presence No ORCID profiles. No Wikidata items. Bio pages without structured data. Critical
Project documentation Major infrastructure, procurement records, impact reports Buried in PDF annual reports. No DOIs. No structured project pages. Critical
Subsidiary network Multiple subsidiaries, joint ventures, strategic partnerships Listed on website as text. No Organization schema with parentOrganization/subOrganization. Moderate
Procurement trails Public/semi-public procurement records, vendor relationships In government procurement databases. Not linked to entity profiles. Moderate

Look at that table. Every single advantage in the left column is real. Verified. Existing. And every single implementation column tells the same story: unstructured, unlinked, invisible to machines.

The authority gap, visualized

I want to make this concrete. The chart below shows the gap between institutional authority potential and actual AI visibility for different organization types. SOEs have the widest gap of any category.

Universities score high on both because academia has spent decades building persistent identifier infrastructure. ORCID, DOI, institutional repositories, ISNI. The tools exist and academics actually use them. Large private corporations score moderately because they invest in SEO and structured data, even if imperfectly. Funded startups, interestingly, sometimes outperform their authority level because they are incentivized to build digital presence from day one.

SOEs sit at the extreme. Highest authority potential. Lowest utilization. That gap is not a problem. It is an opportunity. And a relatively cheap one to close.

What would actually fix this

I am not going to give you a 47-step digital transformation roadmap. That is not how things get done. Here are the specific, concrete actions that would have the highest impact for the lowest effort.

1. ORCID profiles for leadership

Every C-suite executive and board member should have an ORCID. Not because they are academics. Because ORCID is the only globally recognized persistent identifier for people that is free, open, and integrated with every major knowledge system.

When an SOE director publishes a policy paper, speaks at a conference, or is quoted in the Financial Times, that contribution should be linked to a persistent identifier. Right now it is linked to a name string. Name strings are ambiguous. Persistent identifiers are not.

This takes about 15 minutes per person. The return on that 15-minute investment, measured in entity resolution accuracy across AI systems, is enormous.

2. Organization schema on every SOE website

At minimum, every SOE website should have JSON-LD schema declaring: the organization name, founding date, parent organization (the government), subsidiaries, leadership with sameAs links to their ORCID and LinkedIn profiles, official social media accounts, and sameAs links to the Wikidata item, Wikipedia page, Bloomberg profile, and government registry entry.

This is not complex engineering. It is a single block of structured data in the website header. A competent developer implements it in half a day. It connects every fragmented representation of the SOE into a single, authoritative entity.

3. DOIs for major projects and publications

SOEs produce reports, white papers, environmental impact assessments, and technical documentation constantly. None of it has DOIs. All of it should.

Zenodo is free. Registering a DOI takes five minutes. Once a document has a DOI, it exists permanently in the global scholarly infrastructure. AI systems treat DOI-registered content with significantly higher confidence than unregistered PDFs floating on a website [2].

An SOE that registers DOIs for its annual reports, ESG assessments, and major project documentation creates a citable corpus that AI systems can reference with confidence. This is how universities have operated for decades. SOEs just need to adopt the same practice.

4. Structured press pages with article schema

Instead of a news section with dates and headlines, SOE websites should have press pages with proper Article schema. Each press release should have structured data declaring the publisher, author, date, headline, and description. This makes every press release a structured entity event that AI systems can index precisely.

5. Wikidata maintenance

Most SOEs have Wikidata items. Almost none of them are maintained. The leadership is outdated. The financial data is from five years ago. The subsidiary structure is incomplete. Wikidata is the backbone of how AI systems understand entities [3]. An unmaintained Wikidata item is worse than no item at all because it actively misinforms AI systems.

Assign one person to update the Wikidata item quarterly. Revenue figures, leadership changes, subsidiary updates, new office locations. This is not glamorous work. It is foundational.

The cost is embarrassingly low

Let me put numbers on this. For a mid-sized SOE with five C-suite executives and ten major projects per year:

ORCID profiles: 5 people x 15 minutes = 1.25 hours. Cost: zero.

Organization schema: One developer, half a day. Maybe two days if the website is particularly legacy. Cost: whatever you pay a developer for two days.

DOIs for annual publications: 10 documents x 5 minutes = under an hour. Cost: zero on Zenodo.

Wikidata updates: One person, two hours per quarter. Cost: eight hours per year of someone who understands the entity.

Structured press pages: Template change once, then every press release automatically has schema. Cost: half a day of development, once.

Total first-year investment: under one week of developer time plus a few hours of executive onboarding. For an organization that probably spends millions on communications and public affairs annually, this is a rounding error.

The return: accurate AI representation, structured entity presence across all knowledge systems, and a foundation that compounds over time as AI citation becomes more important.

Why this matters now

Governments worldwide are pushing AI adoption in the public sector. The World Bank's leadership training toolkit for SOEs emphasizes digital governance and transparency [4]. National AI strategies increasingly focus on public sector data infrastructure and interoperability.

But all of that policy work assumes the entities themselves are discoverable and accurately represented in the AI systems people actually use. If an SOE is implementing AI internally while remaining invisible to external AI systems, that is a strategic contradiction.

The organizations that structure their entity presence now will be the ones that AI systems cite accurately in 2027, 2028, 2030. The ones that do not will be represented by whatever stale Wikipedia stub and outdated news article happens to rank highest in the training data.

For SOEs specifically, the stakes are higher than for private companies. SOEs represent national capability. When an AI system misrepresents an SOE, it misrepresents the country's industrial capacity. That is not just a corporate communications problem. It is a soft power problem.

A note on procurement and entity infrastructure

I want to address something practical. SOE procurement processes are, by design, bureaucratic. Adding "entity infrastructure" to a procurement scope is not straightforward. It does not fit neatly into existing categories like "IT services" or "communications consulting."

This is actually one of the things I have seen block progress. The work is too small for a major RFP and too specialized for the existing communications vendor. It falls between categories.

The solution is to package it correctly. Entity infrastructure for an SOE is a governance activity, not a marketing activity. It belongs under corporate governance and transparency, not under corporate communications. When framed as "ensuring accurate institutional representation in AI systems and global knowledge infrastructure," it aligns with existing governance mandates around transparency, disclosure, and public accountability [5].

That reframing also helps internally. The governance team understands persistent identifiers and structured data as compliance tools. The communications team sees them as marketing overhead. Frame it right, and the right people own it.

What I have seen work

I have worked with institutional clients on entity infrastructure. Not all SOEs, but organizations with similar structural characteristics: government-linked, heavily regulated, high natural authority, low digital structure.

The pattern that works is small, specific, and cumulative. You do not launch a "digital transformation initiative." You register five ORCID profiles on a Monday morning. You add Organization schema to the website on Tuesday. You update the Wikidata item on Wednesday. By Friday, the entity is more accurately represented in knowledge systems than it has been in its entire history.

The hard part is not the implementation. The hard part is getting the first meeting with someone who has the authority to say "yes, let's do this" and who understands why it matters. Once that person exists, everything else follows quickly.

If you work at an SOE and you are reading this, you are probably that person. Or you know who that person is. The tools are free. The implementation is measured in hours, not months. The only question is whether you understand that AI systems are already representing your organization, badly, whether you participate in that process or not.


Frequently Asked Questions

Why do SOEs score so poorly in AI visibility despite their institutional authority?

Because AI visibility requires structured data, not just authority. SOEs have the authority in abundance: government backing, regulatory filings, decades of public record. But that authority exists in formats AI systems cannot parse efficiently. PDFs, scanned documents, unstructured web pages, JPEG org charts. The authority is real. It is just invisible to machines. Universities do not have this problem because they adopted persistent identifier infrastructure (ORCID, DOI, institutional repositories) decades ago. SOEs never did.

Is ORCID really relevant for SOE leadership who are not academics?

Yes. ORCID is a persistent identifier for people. It was built by academia, but its utility extends to anyone who publishes, speaks at conferences, or contributes to documented projects. SOE directors publish policy papers, present at industry events, and are quoted in media. Linking all of that activity to a persistent identifier allows AI systems to build an accurate, comprehensive profile of that person. Without it, every contribution is an isolated data point that AI systems may or may not connect to the right person. The 15-minute investment in creating an ORCID profile has outsized returns for entity resolution accuracy.

What is the first step an SOE should take to improve AI visibility?

Add Organization schema (JSON-LD) to the website. This is the single highest-impact, lowest-effort action. It declares the organization's name, founding date, leadership, subsidiaries, and sameAs links to external profiles (Wikidata, Wikipedia, Bloomberg, government registry). One block of structured data in the website header connects every fragmented representation into a single authoritative entity. A developer can implement it in half a day.

How does AI misrepresentation of SOEs become a soft power issue?

SOEs represent national industrial and economic capability. When ChatGPT or Gemini provides outdated, inaccurate, or vague information about a country's state-owned enterprises, it shapes how millions of people understand that country's economic capacity. Investment analysts, policy researchers, journalists, and business leaders increasingly use AI systems for preliminary research. If the AI representation of your SOE is a vague paragraph based on a 2019 Wikipedia stub, that is the impression these stakeholders form. Multiply that across every SOE in a country, and you have a systematic underrepresentation of national capability in the global AI information layer.

Can private companies use the same entity infrastructure approach for government-linked work?

Absolutely. Any company that works with government or institutional clients benefits from the same structured approach. Register your leadership on ORCID. Get DOIs for your published reports and project documentation. Add Organization schema with sameAs links. Maintain your Wikidata item if you have one. The entity authority transfer mechanism works in both directions: when you structure your own entity properly, the authority you inherit from institutional clients is better captured and propagated through knowledge systems. I wrote about this mechanism in detail in Why Institutional Clients Make Your Entity Unassailable.

References

  1. Dong, X. et al. "Knowledge Vault: A Web-Scale Approach to Probabilistic Knowledge Fusion." Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2014. Link
  2. Fenner, M. "DOI as a Persistent Identifier for Scholarly Communication." DataCite Blog, 2023. Link
  3. Vrandecic, D. and Krotzsch, M. "Wikidata: A Free Collaborative Knowledgebase." Communications of the ACM, Vol. 57, No. 10, 2014. Link
  4. World Bank Group. "Leadership Training Toolkit for State-Owned Enterprises (SOEs)." Open Knowledge Repository, 2023. Link
  5. Natural Resource Governance Institute. "Precept 6: State-Owned Enterprises." Natural Resource Charter Benchmarking Framework. Link

Linked from

Related notes

2026-03-28

The companies that show up in ChatGPT are the ones that bothered to be verifiable.