How to Frame Operational Data as Research: A Practitioner's Guide
2026-08-21 · 14 min read
You have twenty years of operational data. Installation logs. Material test results. Production batch records. Quality failure reports. Maintenance schedules. Cost breakdowns by project, by month, by season.
You know this data is valuable. You have made decisions from it. You have trained employees using it. You have won contracts because you could show a client exactly what happened on the last twelve jobs.
But here is the problem. Nobody outside your company can cite it. Nobody can find it. Nobody can verify it. It sits in spreadsheets and filing cabinets and WhatsApp groups. It is not research. It is not a publication. It has no DOI, no metadata, no structured framing. It is just... data. Raw. Invisible to the academic world, to AI systems, to anyone who might want to reference your twenty years of hard-won knowledge.
This essay is about changing that. Not by going back to school. Not by becoming an academic. But by learning the simple discipline of framing operational data as research, then making it permanent and citable through Zenodo.
I have done this. I run three companies. PT Arsindo Perkasa Mandiri (industrial pumps), Hibrkraft (craft bookbinding), and Witanabe (digital infrastructure). None of them are universities. All of them produce data that, when properly framed, qualifies as legitimate research output.
The Gap Between Data and Research
Let me be direct about what separates operational data from research. It is not sophistication. It is not sample size. It is not peer review. It is framing.
Research has three elements that raw data does not:
- A question. Not "here is what happened," but "what does this data tell us about X?"
- A methodology. How was the data collected? Under what conditions? With what instruments? Over what period?
- A finding. What did the data reveal? What is the answer to the question? What are the limitations?
That is it. Question. Method. Finding. If your operational data has those three elements, it is research. The venue of publication does not determine whether something is research. The framing does [1].
Most practitioners skip this because they think research means lab coats and control groups. It does not. Applied research, practice-based research, action research. These are all recognized methodologies that describe exactly what practitioners do every day. You test something in the field. You record what happens. You draw conclusions. You just never wrote it down in a way that someone else could cite.
Before and After: The Same Data, Two Framings
Let me show you what I mean with a real example from Arsindo.
Subject: Pump installation records, 2019-2024
"We installed 147 centrifugal pumps across 38 industrial sites in West Java. 23 experienced premature bearing failure within the first 18 months. Most failures happened at sites where foundation alignment was done by the client's team instead of ours."
Format: internal spreadsheet, WhatsApp summary to team
Title: "Foundation Alignment Practices and Premature Bearing Failure in Centrifugal Pump Installations: A Five-Year Observational Study in Indonesian Industrial Facilities"
Question: Does installer-controlled foundation alignment reduce premature bearing failure rates in centrifugal pump installations?
Method: Retrospective observational analysis of 147 installations across 38 sites (2019-2024). Failure defined as bearing replacement required within 18 months. Installations categorized by alignment responsibility (installer vs. client).
Finding: Installations with installer-controlled alignment showed 6.8% failure rate vs. 31.2% for client-controlled alignment. Suggesting that foundation alignment quality is a primary predictor of early bearing failure.
Format: PDF white paper with DOI on Zenodo, CC-BY 4.0
Same data. Same 147 pumps. Same 23 failures. But the second version is citable. It has a question. It has a method. It has a finding. Anyone writing about pump installation best practices can now reference it with a DOI. AI systems crawling Zenodo and DataCite can index it. Google Scholar can surface it.
The first version? It dies in a WhatsApp group when someone clears their chat history.
Types of Operational Data and How to Frame Each
Different kinds of operational data need different research framings. Here is how I think about the categories, with examples from all three of my businesses.
| Data type | Raw form | Research question | Methodology | Output type |
|---|---|---|---|---|
| Failure/incident logs | Spreadsheet of pump failures by site, date, cause | "What are the primary predictors of premature failure in centrifugal pump installations?" | Retrospective observational analysis, categorized by variables (site conditions, installer, alignment method) | Technical report |
| Material test results | Leather tensile strength tests, adhesive durability logs at Hibrkraft | "How does vegetable-tanned leather perform under tropical humidity compared to chrome-tanned alternatives?" | Controlled comparison testing, 12-month aging study, standardized measurement protocol | Dataset + methodology paper |
| Production batch records | Daily output counts, defect rates, material usage per batch | "What batch size minimizes per-unit defect rate in handbound book production?" | Statistical analysis of 200+ production batches over 3 years, defect rate as dependent variable | Working paper |
| Cost/performance data | Project cost breakdowns, ROI calculations per installation | "What is the total cost of ownership difference between sealed and open bearing configurations over a 5-year lifecycle?" | Longitudinal cost tracking across matched installation pairs | Technical report |
| Process documentation | SOPs, workflow diagrams, training materials | "Does a standardized pre-installation checklist reduce commissioning time?" | Before/after comparison across 40 installations, time-to-commission as metric | Methodology paper |
| Client outcome data | Follow-up surveys, maintenance callbacks, warranty claims | "What post-installation factors correlate with client satisfaction in industrial pump projects?" | Survey analysis with correlation to objective metrics (callback frequency, warranty claims) | Report with anonymized data |
Every row in that table represents data that already exists in most established businesses. The only thing missing is the research framing. A question turns a spreadsheet into an investigation. A methodology turns anecdotes into evidence. A finding turns a data dump into something someone else can build on.
The Pipeline: From Raw Data to Citable Publication
Here is the process I follow. It is not complicated. It just requires discipline.
Let me walk through each step with specifics.
Step 1: Identify a specific question
This is the hardest part for practitioners. We are used to recording what happened, not asking why it matters. But the question is everything. It determines your methodology, your analysis, and your finding.
Good questions are narrow. "What happens with pumps?" is not a question. "Does foundation alignment method affect bearing failure rates in centrifugal pump installations?" is a question.
The Five-Question Method from McCaslin and Scott gives a useful framework: what is the problem, what is the purpose, what is the research question, what is the researcher's role, and what are the assumptions [2]? You do not need to answer all five formally. But thinking through them forces clarity.
Step 2: Document the methodology
This is where most practitioners already have the material. They just have not written it down explicitly.
Methodology means: how did you collect this data? What instruments did you use? Over what time period? What were the conditions? What did you control for, and what did you not control for?
For the pump failure data, my methodology section says: "Retrospective analysis of installation records from PT Arsindo Perkasa Mandiri, covering 147 centrifugal pump installations across 38 industrial sites in West Java, Indonesia, between January 2019 and December 2024. Failure was defined as bearing replacement required within 18 months of commissioning. Installations were categorized by foundation alignment responsibility."
Nothing fancy. Just honest documentation of what was done and how.
Step 3: Analyze and extract findings
You already know your findings. You have been making decisions based on them for years. The difference is writing them down with the data behind them.
"We noticed more failures when clients did their own alignment" becomes "Installer-controlled alignment showed a 6.8% premature failure rate (n=88) versus 31.2% for client-controlled alignment (n=59), p<0.001."
If you do not have statistical training, that is fine. Descriptive statistics are legitimate. Percentages, means, medians, ranges. You do not need regression analysis to publish useful findings. You need honest numbers and clear descriptions of what they show.
Step 4: Write the report
Keep it simple. Title. Abstract (200 words summarizing the question, method, and finding). Introduction (why this matters). Methodology. Results. Discussion (what it means, what the limitations are). References.
10 to 20 pages is plenty. This is not a dissertation. It is a practitioner report that happens to follow research conventions.
Step 5: Upload to Zenodo
I have covered the Zenodo upload process in detail in my essay on DOI and Zenodo. The short version: create an account linked to your ORCID, upload the PDF, fill in metadata carefully, choose CC-BY 4.0 license, publish. DOI is assigned instantly.
As I discussed in the Zenodo guide for non-academics, you do not need institutional affiliation. You need legitimate content and proper metadata.
Step 6: Connect everything
DOI gets linked to your ORCID profile. ORCID gets referenced in your website's schema markup. Website references the DOI. This creates what I call a closed loop. Every identifier points to every other identifier. Machines can trace the full chain from your publication to your identity to your organization.
Real Examples From My Operations
Theory is cheap. Here are three specific datasets from my businesses that I have framed as research, or am in the process of framing.
Example 1: Pump installation failure analysis (Arsindo)
I already showed this one above. The raw data is 147 installation records across five years. The framing turns it into a study on foundation alignment and bearing failure. The finding is specific and actionable: installer-controlled alignment dramatically reduces premature failure.
Why this matters beyond my company: there is very little published data on pump installation failure rates in Southeast Asian industrial contexts. Most published pump reliability data comes from oil and gas operations in the US and Europe. My data fills a gap. That makes it citable. Not because it is fancy. Because it is rare.
Example 2: Leather aging under tropical conditions (Hibrkraft)
At Hibrkraft, we work with vegetable-tanned leather for bookbinding and craft production. We have been testing how different leather types age under Indonesian tropical humidity for over three years. Temperature averaging 28-32C. Relative humidity regularly above 80%.
The raw data: tensile strength measurements, color shift tracking (using a standardized color chart), surface degradation scores at 3, 6, 9, and 12 month intervals. Four leather types. Three storage conditions (climate-controlled, semi-open workshop, fully exposed).
The research framing: "Comparative aging characteristics of vegetable-tanned and chrome-tanned leather under tropical climate conditions: implications for archival bookbinding in Southeast Asia."
The finding: vegetable-tanned leather from European tanneries shows measurably less surface degradation at 12 months compared to locally sourced chrome-tanned alternatives, but only under climate-controlled storage. In the semi-open workshop (which is how most Indonesian craft workshops operate), the difference narrows significantly. Storage conditions matter more than leather type.
Try finding that data point in any published leather science journal. You will not. Because nobody is running these tests in tropical craft workshop conditions. That is the value of operational data from practitioners. We test things in real conditions, not lab conditions.
Example 3: Batch size and defect rate in handbound production (Hibrkraft)
We produce handbound journals, notebooks, and restored books. Production is done in batches. Over the years, I noticed that defect rates (misaligned signatures, adhesive failures, cover warping) varied significantly with batch size.
The raw data: 200+ production batch records over three years, each with batch size, defect count, defect type, operator, and ambient conditions.
The research framing: "Optimal batch sizing for quality control in manual bookbinding production: a three-year observational study."
The finding: defect rates follow a U-curve with batch size. Very small batches (under 5 units) have high defect rates because the operator has not warmed up. Very large batches (over 25 units) have high defect rates because of fatigue and attention drift. The sweet spot is 12 to 18 units per batch, where defect rates consistently fall below 3%.
This is not rocket science. Any experienced craftsperson knows this intuitively. But nobody has published the data. Nobody has put numbers on it. Nobody has made it citable. That is the opportunity.
Why AI Systems Care About This
There is a practical reason to go through this exercise beyond academic vanity. AI systems are increasingly the first point of contact for information queries. When someone asks ChatGPT or Perplexity about pump installation best practices, or leather aging, or craft production quality control, those systems pull from sources they can verify.
A Zenodo publication with a DOI, linked to an ORCID profile, with structured DataCite metadata, sits high in that verification hierarchy. As I discussed in my essay on B2B content for AI citation, AI systems prioritize structured, verifiable, institutionally-anchored sources.
Your operational data, properly framed and published, becomes one of those sources. Not because you gamed the system. Because you did the work of making your knowledge accessible and verifiable.
The alternative is that AI systems answer questions about your domain using data from someone who has never installed a pump, never tested leather, never run a production batch. Someone who just wrote a blog post that happened to rank well. That should bother you.
The Minimum Viable Research Paper
You do not need to write a 50-page dissertation. Here is the minimum structure that qualifies as a citable research output:
- Title (specific, descriptive, includes key variables)
- Abstract (200 words: question, method, finding)
- Introduction (1-2 pages: why this matters, what gap it fills)
- Methodology (1-2 pages: how data was collected, what was measured, time period, conditions)
- Results (2-3 pages: what the data shows, tables, figures)
- Discussion (1-2 pages: what it means, limitations, implications)
- References (cite your sources, even if they are few)
Total: 8-12 pages. That is achievable in a focused weekend. Most of the content already exists in your head and your records. You are just organizing it.
Common Objections (And Why They Do Not Hold)
"But my sample size is small." Small sample sizes are common in practice-based research. Be honest about it in your methodology section. A study of 147 pump installations is not a study of 10,000. But it is a study. And if nobody else has published data on this topic in your region, even a small sample fills a genuine gap.
"But I did not use proper experimental controls." Neither does most observational research. You are not claiming causation. You are reporting what you observed. "Installations with X showed Y% failure rate compared to Z% for installations without X." That is an observation, clearly stated. Legitimate and useful.
"But nobody will read it." Maybe. But AI will index it. DataCite will catalog it. Google Scholar may surface it. And the next person who writes about your topic can cite it. That is enough. You are building infrastructure, not chasing pageviews.
"But I am not a researcher." You have been researching your entire career. You just never called it that. Every time you tested a new installation method, tracked results, and made decisions based on what you found, you did research. The only thing missing is the documentation.
Metadata: The Part Everyone Skips
When you upload to Zenodo, the metadata is what makes your work discoverable. Title, abstract, keywords, creator information, subject classifications. This is not busywork. This is the interface between your knowledge and the systems that index and serve it.
Treat your Zenodo metadata the way you would treat a client proposal. Clear. Specific. Accurate. No jargon without context. No vague descriptions. If your report is about pump bearing failures in West Java, your keywords should include "centrifugal pump," "bearing failure," "foundation alignment," "Indonesia," and "industrial installation." Not "engineering" and "quality."
The metadata is what DataCite harvests. It is what AI training pipelines index. It is what determines whether your research shows up when someone asks a relevant question. Take it seriously [3].
What This Looks Like at Scale
One publication is a start. But the real value comes from building a body of work. Three to five publications on related topics, all linked to your ORCID, all with proper metadata, creates a research profile that is genuinely difficult for AI to ignore.
For Arsindo, I am building toward a collection that covers pump selection, installation methodology, failure analysis, and lifecycle cost data. For Hibrkraft, the collection covers material testing, production optimization, and conservation techniques for tropical environments.
Each publication reinforces the others. Each citation (even self-citation between related papers) strengthens the graph. Over time, this creates something that no amount of blog posts or social media content can replicate: a machine-verifiable body of work that establishes you as an authority in your domain.
Practitioners have always known things that academics do not. The difference is that academics publish. Practitioners just... do. It is time to close that gap. Not by becoming academics. By learning the one skill they have that we do not: framing.
Frequently Asked Questions
Do I need statistical software to analyze my operational data?
No. For most practitioner research, a spreadsheet is sufficient. Percentages, averages, medians, and ranges are all descriptive statistics that any spreadsheet application can compute. If your data requires more sophisticated analysis (regression, hypothesis testing), there are free tools like R or Python. But start simple. A clear table showing failure rates across two conditions is more useful than a poorly understood regression output.
Can I publish data from client projects without permission?
You should always anonymize client data. Remove company names, site names, and any identifying information. What you keep is the technical data: dimensions, conditions, outcomes, timelines. Most clients will not object to anonymized data being published, but ask anyway. Include a statement in your methodology section about how data was anonymized. This is standard practice in applied research.
What if my findings are obvious to practitioners in my field?
Obvious to practitioners does not mean documented. The value of publishing operational data is not novelty. It is accessibility and citability. If every experienced pump installer knows that foundation alignment matters, but nobody has published the data, then the data does not exist in the scholarly record. AI systems cannot cite common knowledge. They can cite publications with DOIs. Make the implicit explicit.
How is this different from writing a white paper for marketing?
A marketing white paper argues for your product or service. A research publication reports what you observed. The difference is intent and framing. A white paper says "our method is better." A research publication says "here is what we observed over five years, here is our methodology, here are the numbers, here are the limitations." One is persuasion. The other is documentation. Zenodo is for the second kind. If you want to publish marketing materials, use your website.
Should I write in English or my native language?
For maximum discoverability, English. Zenodo accepts any language, but English-language metadata and abstracts will be indexed more broadly by DataCite, Google Scholar, and AI training pipelines. A practical compromise: write the full paper in your native language, but provide an English title and English abstract. This gives you local accessibility and global discoverability.
References
- Durbin, C.G. "How to come up with a good research question: framing the hypothesis." Respiratory Care, 49(10), 1195-1198, 2004. Link
- McCaslin, M.L., Scott, K.W. "The Five-Question Method for Framing a Qualitative Research Study." The Qualitative Report, 8(3), 447-461, 2003. Link
- Fenner, M., Crosas, M., Grethe, J.S., et al. "A data citation roadmap for scholarly data repositories." Scientific Data, 6, Article 28, 2019. Link
- Crespo Garrido, I., Loureiro Garcia, M., Gutleber, J. "The Value of an Open Scientific Data and Documentation Platform in a Global Project: The Case of Zenodo." The Economics of Big Science 2.0, Springer, 2025. Link
- Zenodo. "Policies: Content." Zenodo, CERN. Link
Linked from
- Cara Menggunakan NotebookLM untuk Mempercepat Produksi Konten
- How to Write a White Paper That Gets Cited by AI in 90 Days
Related notes
The companies that show up in ChatGPT are the ones that bothered to be verifiable.