Indonesia publishes a lot of research. The numbers are not small. Scopus indexes tens of thousands of Indonesian papers every year. The country has over 4,700 higher education institutions. Some of them, like Universitas Indonesia and Institut Teknologi Bandung, rank respectably in regional tables.

And yet. Ask ChatGPT about a research topic where Indonesian universities have published extensively, and the answer will cite Singapore, Australia, maybe Thailand. Not Indonesia.

This is not about quality. This is about infrastructure.

Indonesian universities have a research output problem that is actually a visibility problem. They publish, but they do not connect. They produce knowledge, but that knowledge does not propagate into the systems that AI uses to decide who is credible.

The gap is real. And it is widening.

Key concept: Publication volume does not equal AI visibility. AI systems cite entities they can verify across multiple structured databases. Indonesian universities publish into Scopus but fail to build the entity infrastructure that connects their researchers, departments, and institutions to the knowledge graph layer AI depends on.

The numbers tell one story. AI tells another.

Let me show you something. Southeast Asian countries publish research at vastly different rates. But the relationship between publication volume and AI citation is not linear. It is not even close.

Singapore publishes fewer papers per capita than Indonesia does in total volume. But per-publication citation impact in Singapore is among the highest in the world. Malaysia has the largest publication volume in ASEAN. Thailand shows stronger citation impact per paper than Indonesia despite publishing less.

The ASEAN Higher Education Report (2025) from Times Higher Education and Elsevier found that Singapore's share of publications in the top 10% most cited globally is exceptional. Malaysia's robust output is now coupled with improving citation rates. Indonesia's volume is large but its per-paper impact lags behind [1].

Here is the picture in chart form. These are approximate AI citation rates, meaning how often universities from each country appear as sources in AI-generated answers on research topics, compared against their raw Scopus publication output.

Look at the gap between the gray bars (publication volume) and the tan bars (AI citation). Singapore has a small output but dominates AI answers. Indonesia has a large output and barely registers.

That gap is entity infrastructure. Or rather, the absence of it.

What universities have vs. what they are missing

Indonesian universities are not starting from zero. They have institutional capacity. They have researchers. They have output. What they lack is the connective tissue that makes that output machine-readable, cross-referenced, and verifiable at the entity level.

Infrastructure layer What most universities have What they are missing
Publication output Scopus-indexed papers, SINTA rankings DOI coverage is inconsistent. Many local journals lack DOIs entirely.
Researcher identity Google Scholar profiles (often incomplete) Low ORCID adoption. Name disambiguation is a known problem for Indonesian surnames.
Institutional identity Official websites, some Wikidata entries No consistent ROR (Research Organization Registry) linking. Inconsistent schema markup.
Data repositories Institutional repositories (often closed or poorly indexed) Almost no Zenodo, Figshare, or Dataverse deposits with proper metadata.
Entity connections Internal citation networks within Indonesian journals No structured links between researchers, institutions, grants, and outputs. No sameAs properties.
Web presence University websites with faculty listings Faculty pages lack JSON-LD, ORCID links, or structured publication lists. Many are PDF-only CVs.
Wikipedia / Wikidata Top 10 universities have Wikipedia articles Hundreds of institutions have no Wikidata item. Researcher items are almost nonexistent.

Every row in that "missing" column is something that AI systems use to evaluate credibility. When an LLM encounters a claim and needs to attribute it, it looks for entities it can verify. Not just "a paper exists," but "this person at this institution published this work, which is cross-referenced in these databases, with this persistent identifier."

If your researcher has an ORCID that connects to Crossref that connects to their institutional page that connects to Wikidata, that is a verification chain. AI can follow it.

If your researcher has a Google Scholar profile with a misspelled name and no institutional email, that chain is broken. AI walks away.

The ORCID problem is bigger than you think

I have written before about how ORCID is not just for academics. But for universities specifically, the adoption gap is devastating.

ORCID adoption varies wildly across countries. Research from Firdaus et al. (2022) specifically identified Indonesian authors as having significant disambiguation problems due to common surnames, transliteration inconsistencies, and metadata deficiencies [2]. This is not a minor technical issue. It means AI cannot reliably connect Indonesian researcher A to their publications, their grants, their institution.

In France, a national survey found uneven ORCID adoption with higher uptake in natural sciences [3]. In Spain, larger universities showed better integration than smaller ones. Indonesia? The data is harder to find, which is itself part of the problem. When your adoption rate is so low that nobody bothers to study it, that tells you something.

Compare this to Singapore. NUS, NTU, and SMU have institutional ORCID integration built into their faculty management systems. When a researcher at NUS publishes a paper, their ORCID is automatically updated. The chain is unbroken. AI can trace it.

In Indonesia, a researcher publishes, the paper gets a Scopus entry, and then... nothing. No ORCID update. No Crossref enrichment. No institutional profile that links to the paper. The knowledge exists. But it exists in isolation.

Why Webometrics obsession makes this worse

Here is an irony that hurts. Indonesian universities are obsessed with Webometrics rankings. A recent analysis found that Indonesia leads the world in press coverage of Webometrics results, with 94 publications celebrating ranking changes, more than any other country [4].

Webometrics measures web presence. University websites, open access content, Google Scholar visibility. It is a reasonable metric. But it has become a vanity metric in Indonesia. Universities celebrate moving up five spots. Press offices write stories. Nobody asks whether that web presence is structured in a way that AI systems can actually use.

You can have the best Webometrics score in Southeast Asia and still be invisible to ChatGPT. Because Webometrics measures quantity of web presence. AI systems need quality of entity infrastructure. These are different things.

A university that publishes 10,000 pages on its website but has no JSON-LD, no ROR identifier in its schema, no ORCID integration for faculty, and no Wikidata items for its departments is just noise to an LLM. Lots of text. No structure. No verification chain.

The training data problem

This connects directly to how AI training data decides who gets cited. LLMs learn entity credibility from the training corpus. The training corpus is built from sources like Common Crawl, Wikipedia, arXiv, PubMed, and curated datasets.

When Indonesian university research appears in Scopus but not in Wikipedia, not in well-structured institutional repositories, not in open data platforms with proper metadata, it exists in exactly one place in the training pipeline. One signal is not enough for an LLM to build entity confidence.

Singapore's NUS appears in:

  • Wikipedia (detailed article, hundreds of citations)
  • Wikidata (structured properties, sameAs links)
  • ROR (institutional identifier)
  • Crossref (as publisher and affiliation)
  • ORCID (institutional member with automated linking)
  • Multiple open data repositories
  • Institutional schema markup on every faculty page

A mid-tier Indonesian university appears in:

  • Scopus (publication records)
  • SINTA (Indonesian national index)
  • Maybe a brief Wikipedia stub
  • Its own website (unstructured)

Four signals vs. seven or more. And not just quantity. The quality of those signals matters. Wikipedia and Wikidata are in every major training dataset. SINTA is not.

The institutional repository failure

Most Indonesian universities have institutional repositories. This sounds good. It is not.

A 2024 systematic review of institutional repository challenges found that common problems include poor metadata, low discoverability, lack of persistent identifiers, and inadequate interoperability with global systems [5]. Indonesian repositories suffer from all of these.

Many are running DSpace or EPrints installations from a decade ago. Metadata is inconsistent. DOIs are not minted for locally published work. OAI-PMH harvesting is broken or not configured. The result is that these repositories are invisible to Common Crawl, invisible to Google Scholar in many cases, and completely invisible to the curated datasets that feed LLM training.

This is infrastructure work. It is not glamorous. Nobody gets a promotion for fixing OAI-PMH endpoints. But it is the difference between existing and being findable.

What the fix looks like

This is not a technology problem. The technology exists. ORCID is free. DOIs can be minted through Zenodo at no cost. Wikidata is open for editing. JSON-LD is a few lines of code. ROR identifiers are free to claim.

The problem is institutional. It is a question of who owns this work. In most Indonesian universities, nobody does. The library manages the repository. The IT department manages the website. The research office manages Scopus reporting. Nobody manages the entity layer that connects all three.

Here is a practical starting point for any Indonesian university that wants to stop being invisible to AI:

Month 1-2: Identity layer. Ensure every active researcher has an ORCID. Not optional. Not "encouraged." Required for any publication submission. Connect ORCID to the institutional email domain. This alone solves the disambiguation problem.

Month 3-4: Institutional identity. Claim or verify your ROR entry. Add Organization schema markup to your homepage with sameAs properties pointing to Wikidata, ROR, ISNI. Create or improve your Wikidata item with proper claims and references.

Month 5-6: Repository upgrade. Audit your institutional repository. Fix broken metadata. Enable OAI-PMH. Mint DOIs for theses and locally published papers through Crossref or DataCite. Deposit research data in Zenodo or Figshare with proper metadata linking back to the institution.

Month 7-8: Faculty pages. Add JSON-LD Person schema to every faculty profile page. Include ORCID, affiliation, notable publications. Link to the institutional Organization entity. This gives AI systems a structured entry point.

Month 9-12: Connection layer. Build the links between all of these. Researcher ORCIDs should point to institutional profiles. Institutional profiles should point to Wikidata. Publications should link to researchers and institutions through DOIs and persistent identifiers. Create a verification chain that AI can follow from any starting point.

None of this requires new technology. It requires someone to decide it matters.

Why this matters beyond vanity

This is not about appearing in ChatGPT for ego reasons. The visibility gap has material consequences.

International research collaboration increasingly depends on discoverability. Funding bodies use AI-assisted literature reviews. Industry partnerships start with "what does the research say?" And if the research is invisible, the partnership goes to the university that is visible.

Indonesia has the fourth largest population in the world. It has thousands of researchers doing work that matters. Climate science. Tropical medicine. Islamic finance. Biodiversity. These are areas where Indonesian institutions should be definitive sources.

They are not. Because the entity infrastructure does not exist.

The fix is not expensive. It is not technically difficult. It requires institutional will and someone who understands that the game has changed. Publication is no longer the finish line. Structured, verifiable, machine-readable presence across multiple databases is.

The universities that figure this out in the next two years will become the ones AI cites. The rest will keep publishing into silence.


Frequently Asked Questions

Why are Indonesian universities invisible to AI despite high publication volume?

Publication volume alone does not create AI visibility. AI systems rely on entity verification across multiple structured databases: Wikipedia, Wikidata, ORCID, Crossref, ROR, and institutional schema markup. Indonesian universities publish into Scopus and SINTA but rarely connect that output to the broader entity infrastructure that LLMs use to assess credibility. The chain of verification is broken at multiple points.

Is this problem specific to Indonesia or common across Southeast Asia?

The problem exists across much of the developing world, but Indonesia's scale makes it particularly striking. With over 4,700 higher education institutions and significant research output, the gap between volume and AI visibility is larger than in neighboring countries. Singapore and Malaysia have invested more heavily in entity infrastructure at the institutional level. Thailand shows stronger per-paper citation impact. Indonesia's gap is the largest relative to its potential.

How much does it cost to build university entity infrastructure?

Almost nothing in direct costs. ORCID membership for institutions is affordable. ROR identifiers are free. Wikidata editing is free. Zenodo deposits are free. JSON-LD markup is a few hours of developer time. The real cost is institutional coordination: getting the library, IT, research office, and faculty to work together on something that nobody currently owns.

Can improving Webometrics rankings help with AI visibility?

Not directly. Webometrics measures web presence quantity, including things like number of pages indexed and Google Scholar citations. AI visibility requires structured entity infrastructure: persistent identifiers, schema markup, cross-database linking. A university can have an excellent Webometrics score and still be invisible to ChatGPT because the web presence is unstructured text rather than machine-readable entities.

What should a university prioritize first: ORCID, Wikidata, or schema markup?

ORCID. Researcher identity is the foundation. Without unique, persistent identifiers for your researchers, nothing else connects properly. ORCID creates the anchor that links publications to people to institutions. After ORCID adoption reaches a critical mass, move to institutional identity (ROR, Wikidata) and then to schema markup on faculty pages. The order matters because each layer depends on the one before it.

References

  1. Times Higher Education and Elsevier. "How Universities Are Shaping ASEAN's Tomorrow." ASEAN Report 2025. Link
  2. Firdaus et al. "Identification of Indonesian Authors Using Deep Neural Networks." Computer Engineering and Applications Journal, vol. 11, 2022, pp. 15-24. Link
  3. Bouchard and Boudry. "Knowledge and Use of the ORCID Author Identifier in France: A National Survey." Learned Publishing, vol. 38, 2025. Link
  4. Isidro Aguillo. "Predatory University Rankings Jeopardise the Value of Webometrics." LSE Impact Blog, March 2026. Link
  5. Asistdl. "Current Challenges and Future Directions for Institutional Repositories." Journal of the Association for Information Science and Technology, 2025. Link

Related notes

2026-03-28

The companies that show up in ChatGPT are the ones that bothered to be verifiable.