Save products you love by clicking the heart icon.
Data source: graph-research corpus (17,553 papers, 100% taxonomy saturation) · Generated: 2026-08-06
The ISO GQL standard — the first international standard for property-graph querying — has triggered a research surge unlike anything else in the graph field. This analysis of 309 papers in the graph-query-languages taxonomy category shows the field doubling in 2026: 119 papers in the first eight months, and +235.7% year-over-year growth (141 papers in the last 12 months vs 42 in the prior 12). We dissect the surge into four research threads — (1) formal theory catching up to the standard, (2) a natural-language-to-GQL (NL2GQL) tooling wave with public benchmarks, (3) engine, path, and schema infrastructure, and (4) GQL reused as a primitive for adjacent problems such as ontology-mediated queries and cyber-provenance analysis. We map the author community (dominated by a small committee-adjacent group plus a China-based NL2GQL cluster) and identify the field's sharpest white space: only 8 review/survey papers against 309 total — the strongest editorial and research gap in the corpus.
For two decades, graph querying was fragmented: Cypher, SPARQL, Gremlin, PGQL, GSQL and proprietary dialects each owned a slice of the market. The ISO/IEC 39075:2024 GQL standard changed the calculus — a single international standard for property graphs, with SQL/PGQ as its relational companion. What the corpus reveals is that the research community has now fully engaged: not just evaluating the standard, but attacking its semantics, building tooling on top of it, and reusing it as a substrate for other problems.
This report is a data-driven analysis of that surge: what is being published, in which aspects, by whom, and where the gaps are.
Corpus. The analysis uses the graph-research corpus (17,553 papers, 100% of 160 taxonomy cells filled, spanning 2021–2026). Papers are auto-classified into a 20×8 taxonomy (20 categories × 8 aspects: theory, mechanism, method, application, development, systems, evaluation, review).
Selection. The core population is the graph-query-languages category (309 papers). A secondary scan targets explicit standard mentions: "GQL" (37 papers), "Cypher" (84), "SPARQL" (192). All counts are derived from titles and abstracts.
Metrics. We compute (a) year-over-year growth from two non-overlapping 12-month windows, (b) aspect-level trajectories, (c) keyword co-occurrence, and (d) author and venue distributions. Growth is windowed to avoid partial-year bias.
| Year | Category papers | Δ YoY |
|---|---|---|
| 2021 | 8 | — |
| 2022 | 44 | +450% |
| 2023 | 39 | −11% |
| 2024 | 49 | +26% |
| 2025 | 50 | +2% |
| 2026 (8 mo) | 119 | +238% run-rate |
The last 12 months contributed 141 papers versus 42 in the prior 12 months — a +235.7% growth rate that makes graph query languages the fastest-accelerating category in the entire corpus (ahead of GraphRAG at +130%, Graph Databases at +149%, and Ontologies at +133%).
The 2026 aspect mix shows the nature of the surge:
| Aspect | 2026 papers | Share |
|---|---|---|
| Application | 31 | 26% |
| Method | 28 | 24% |
| Systems | 23 | 19% |
| Theory | 11 | 9% |
| Mechanism | 10 | 8% |
| Evaluation | 9 | 8% |
| Review | 4 | 3% |
| Development | 3 | 3% |
Two-thirds of 2026 output is Application + Method + Systems — the field is past the "what is GQL?" stage and into "use it, optimise it, automate against it".
The 2024–2026 period is closing the gap flagged by the standards community's own assessment that "rapid industrial development left the academic community trailing." Key results:
The largest recent application/method cluster is natural-language-to-GQL:
Implication: property graphs are increasingly queried in natural language, and the tooling is now benchmark-graded.
GQL is being reused beyond graph databases:
The category's author distribution is concentrated:
| Author | Papers |
|---|---|
| Mahmoud Abo Khamis | 5 |
| Diego Figueira | 5 |
| Makbule Gulcin Ozsoy | 3 |
| Lukas Arzoumanidis | 3 |
| Subhasis Dasgupta | 3 |
| Ahmet Kara | 3 |
| Yutong Ye | 3 |
| Rémi Morvan | 3 |
Within explicit GQL papers, the theory cluster (Angela Bonifati, Renzo Angles, Nadime Francis, Wim Martens, Amélie Gheerbrant, Leonid Libkin, Liat Peterfreund) dominates, alongside a China-based NL2GQL cluster (Yuanyuan Liang, Tingyu Xie). The field is driven by a small number of groups — which is an opportunity: most practitioners have no mental model of GQL yet.
The category is 55% arXiv (169/309), reflecting both the theory-heavy preprint culture and the conference-to-preprint pipeline of the database community.
The sharpest signal is the review gap:
| Aspect | Papers |
|---|---|
| Review | 8 |
| Development | 14 |
| Mechanism | 25 |
| Evaluation | 28 |
Only 8 review/survey papers against 309 total — nobody has yet written the definitive practitioner "state of GQL" survey. This is the strongest white-space cell in the category and one of the strongest in the whole corpus.
Secondary gaps:
Suggested roadmap: (short-term) practitioner survey + compositionality engineering analysis; (medium-term) dialect performance benchmarks, NL2GQL robustness under adversarial/ambiguous queries; (long-term) GQL as a uniform substrate for ontology-mediated and provenance querying.
GQL is not a peripheral database-specification story. It is the highest-momentum research subject in the corpus in 2026: a mature standard, a theory community that has caught up, an NL2GQL tooling wave with public benchmarks, and demonstrable applications in security and ontology-mediated querying. The field's own white space — a missing survey, unstudied compositionality implications, and the agentic convergence — offers unusually strong opportunities for both researchers and practitioners.
This analysis was produced from the graph-research corpus (github.com/tobias-weiss-ai-xr/graph-research): 17,553 papers, 100% taxonomy saturation, 2021–2026. Full methodology and raw statistics available in statistics.json and the graph-query-languages category.