Back openDesk Edu for a sovereign, open-source education — every vote counts.
Vote nowSave products you love by clicking the heart icon.
Data source: graph-research corpus (26,496 papers, 88.5% taxonomy saturation) · Generated: 2026-09-04 · Initial edition: 2026-08-06
The ISO GQL standard — the first international standard for property-graph querying — has triggered a research surge unlike anything else in the graph field. This analysis of 584 papers in the graph-query-languages taxonomy category shows the field compounding through 2026: 243 papers in the first eight months alone, and +154.4% year-over-year growth (290 papers in the last 12 months vs 114 in the prior 12) — still the fastest-accelerating category in the entire 26,496-paper corpus. We dissect the surge into four research threads — (1) formal theory catching up to the standard, (2) a natural-language-to-GQL (NL2GQL) tooling wave with public benchmarks, (3) engine, path, and schema infrastructure, and (4) GQL reused as a primitive for adjacent problems such as ontology-mediated queries and cyber-provenance analysis. We map the author community (dominated by a small committee-adjacent group plus a China-based NL2GQL cluster) and probe the field's white space: only 15 review/survey papers against 584 total — the smallest cell in the category, though a hierarchical Bayesian cross-check shows it sits exactly on the corpus-wide pattern of young fields; the genuinely missing artefact is the practitioner survey, and the field's standout surplus is benchmark-grade evaluation work.
For two decades, graph querying was fragmented: Cypher, SPARQL, Gremlin, PGQL, GSQL and proprietary dialects each owned a slice of the market. The ISO/IEC 39075:2024 GQL standard changed the calculus — a single international standard for property graphs, with SQL/PGQ as its relational companion. What the corpus reveals is that the research community has now fully engaged: not just evaluating the standard, but attacking its semantics, building tooling on top of it, and reusing it as a substrate for other problems.
This report is a data-driven analysis of that surge: what is being published, in which aspects, by whom, and where the gaps are.
Corpus. The analysis uses the graph-research corpus (26,496 papers, 177 of 200 taxonomy cells filled = 88.5% saturation, spanning 2011–2026). Papers are auto-classified into a 20×10 taxonomy (20 categories × 10 subcategories: theory, mechanism, method, application, development, systems, evaluation, review, survey, experiment).
What changed since the August edition. Between 2026-08-06 and 2026-09-04 the corpus grew from 17,553 to 26,496 papers (+51%) via a full six-host fleet re-fetch (arXiv + OpenAlex, 36-month window, de-duplicated). The category grew from 309 to 584 papers. The re-fetch also broadened the category definition: keyword-adjacent publications (ontology engineering, semantic-web tooling, RDF benchmarks) now make up roughly half of the category. We therefore report the official pipeline numbers throughout, plus a GQL-core cross-check (286 papers whose title/abstract contain both a graph term and a query term): its trajectory is materially slower but still steep — +98.4% YoY (125 vs 63), 97 core papers in 2026. Treat 584 as the category's outer bound and 286 as its inner bound.
Selection. The core population is the graph-query-languages category (584 papers). A secondary scan targets explicit standard mentions within the category (title + abstract, word-boundary match): "GQL" (16 papers), "Cypher" (37), "SPARQL" (231). All counts are derived from titles and abstracts.
Metrics. We compute (a) year-over-year growth from two non-overlapping 12-month windows, (b) aspect-level trajectories, (c) keyword co-occurrence, and (d) author and venue distributions. Growth is windowed to avoid partial-year bias.
| Year | Category papers | Δ YoY |
|---|---|---|
| 2021 | 8 | — |
| 2022 | 46 | +475% |
| 2023 | 47 | +2% |
| 2024 | 97 | +106% |
| 2025 | 131 | +35% |
| 2026 (8 mo) | 243 | +178% run-rate |
The last 12 months (Sep 2025 – Aug 2026) contributed 290 papers versus 114 in the prior 12 months — a +154.4% growth rate that keeps graph query languages the fastest-accelerating category in the entire corpus (ahead of GraphRAG at +86.3%, Graph Analytics at +84.7%, and Ontologies and Graph Databases at +63.8% each).
The 2026 aspect mix shows the nature of the surge:
| Aspect | 2026 papers | Share |
|---|---|---|
| Method | 84 | 35% |
| Application | 38 | 16% |
| Evaluation | 34 | 14% |
| Theory | 34 | 14% |
| Systems | 30 | 12% |
| Mechanism | 13 | 5% |
| Review | 7 | 3% |
| Development | 3 | 1% |
Method now leads by a wide margin — the field's energy has shifted from "what is GQL?" to "automate against it": NL2GQL generation, translation methods, and benchmark papers dominate 2026 output, with a solid evaluation tier (34 papers, up from 9 in the August edition's count) grading them.
The 2024–2026 period is closing the gap flagged by the standards community's own assessment that "rapid industrial development left the academic community trailing." Key results:
The largest recent application/method cluster is natural-language-to-GQL:
Implication: property graphs are increasingly queried in natural language, and the tooling is now benchmark-graded.
GQL is being reused beyond graph databases:
The category's author distribution is concentrated:
| Author | Papers |
|---|---|
| Mahmoud Abo Khamis | 5 |
| Diego Figueira | 5 |
| Leonid Libkin | 4 |
| Wim Martens | 4 |
| Ginwa Fakih | 4 |
| Yannis Tzitzikas | 4 |
| Lukas Arzoumanidis | 3 |
| Ruben Taelman | 3 |
The widened corpus net pulls in semantic-web adjacent authors (Tzitzikas, Taelman, Arcangelo Massari) — another reflection of the broader category definition. Within explicit GQL papers, the theory cluster (Angela Bonifati, Renzo Angles, Nadime Francis, Wim Martens, Amélie Gheerbrant, Leonid Libkin, Liat Peterfreund) dominates, alongside a China-based NL2GQL cluster (Yuanyuan Liang, Tingyu Xie). The field is driven by a small number of groups — which is an opportunity: most practitioners have no mental model of GQL yet.
The category is now 34% arXiv (201/584) — down from 55% in the August edition, because the fleet re-fetch added thousands of DOI/OpenAlex-only records. The theory-heavy preprint culture remains visible, but the category's center of gravity has shifted toward the published record.
The sharpest signal is the review gap:
| Aspect | Papers |
|---|---|
| Review | 15 |
| Development | 16 |
| Mechanism | 33 |
| Evaluation | 94 |
Only 15 review/survey papers against 584 total (2.6%) — and in the GQL-core subset it is 4 of 286 (1.4%). Nobody has yet written the definitive practitioner "state of GQL" survey.
One caveat from our own hierarchical Bayesian cross-check (empirical-Bayes Gamma shrinkage across all 200 corpus taxonomy cells, graph-research/docs/bayesian-gap-analysis.md): the review cell sits exactly on the corpus-wide pattern for its expected size (n = 15 vs 15.0 expected; posterior rate ratio 1.0, 95% CI 0.58–1.54, P(gap) = 0.008). Thin review cells are a global property of young fields — so the honest claim is not "GQL's review cell is an extreme corpus anomaly" but "the survey desert of a young field". Two sharper Bayesian signals for GQL: its evaluation cell is significantly over-represented (94 papers vs 38.2 expected, rate ratio 2.41, P > 0.95) — a benchmarked field by structure — and its only notable deficit is the survey cell (0 of 5.1 expected, P(gap) = 0.93). For calibration, the corpus cells with ≥ 95% posterior gap probability live elsewhere: applications/method (141/422.9), graph-construction/theory (30/164.2), graph-theory/application (86/343.5), and five empty survey cells.
Secondary gaps:
Suggested roadmap: (short-term) practitioner survey + compositionality engineering analysis; (medium-term) dialect performance benchmarks, NL2GQL robustness under adversarial/ambiguous queries; (long-term) GQL as a uniform substrate for ontology-mediated and provenance querying.
GQL is not a peripheral database-specification story. It is the highest-momentum research subject in the corpus in 2026: a mature standard, a theory community that has caught up, an NL2GQL tooling wave with public benchmarks, and demonstrable applications in security and ontology-mediated querying. Even on the conservative GQL-core subset the field is growing at ~98% year-over-year. The field's own white space — a missing survey, unstudied compositionality implications, and the agentic convergence — offers unusually strong opportunities for both researchers and practitioners.
This analysis was produced from the graph-research corpus (github.com/tobias-weiss-ai-xr/graph-research): 26,496 papers, 88.5% taxonomy saturation (177/200 cells), 2011–2026, generated 2026-09-04. Full methodology and raw statistics available in statistics.json and the graph-query-languages category.