CATS QDA
A local-first workspace where qualitative researchers analyze sensitive interview data with AI assistance they can inspect, override, and audit.
Built on our work in AI for qualitative research: Using Large Language Models and Generative AI to Scale Qualitative Data Analysis.
The problem
Rigorous qualitative coding, memoing, and theme development take substantial researcher time, so valuable interview and document collections are often analyzed only at limited scale, and the usual AI shortcuts hide exactly the judgments researchers need to see.
An end-to-end analysis workspace that runs entirely on the researcher’s own computer: transcripts, codebooks, themes, memos, search, and data-grounded chat, with agentic workflows that propose codes and organize themes while recording an execution trace the researcher can review stage by stage.
What it looks like
Research basis
The workbench builds the GATOS codebook-generation method, published in Humanities and Social Sciences Communications and applied at scale in the Journal of Engineering Education, into an interactive system. The pipeline a researcher watches run (corpus understanding, candidate generation, novelty analysis, theme organization, quality review) is the published method, made inspectable.
What it does today
- Organizes projects around research questions, with transcript ingestion, speaker attribution, dialogue acts, filtering, and version controls.
- Answers researcher questions from the corpus with grounded chat that cites source turns and reports which retrieval tools ran.
- Generates candidate codebooks through a multi-stage agent pipeline and keeps a persistent execution trace of every stage.
- Structures codebooks with definitions, abstraction levels, theme grouping, and explicit acceptance state: the researcher accepts or rejects, not the model.
- Links analytic memos to files, codes, and themes with provenance labels.
- Runs local-first, with the database, vector search, and language models on the researcher’s own machine, so sensitive corpora never leave it.
What it does not do
- Not hosted and not installable without a developer-oriented local stack; explicitly single-user and loopback-only, with no authentication or multi-user isolation.
- No completed external pilot or user study.
- The quality indicators it reports (saturation, coverage, coherence, balance) have not been validated against benchmarks.
What it does not show yet
- A working demonstration against synthetic data is evidence that the system exists and runs, not that it improves research quality, productivity, or coding validity. Those claims wait on the planned pilot and evaluation.
- Agent-proposed codes and themes are candidates for researcher judgment, not findings. The design thesis is inspectability, and the honest corollary is that the human review it enables is still required.
Outcomes
- Turns the lab’s published GATOS method into an interactive, inspectable platform.
- Demonstrates full-stack research engineering: durable agent execution, provenance records, and grounded retrieval over a private corpus.
Limitations and responsible use
- Use with real research data stays inside the documented local trust boundary and applicable IRB and data-use agreements; the platform makes analysis local, not consent unnecessary.
- All interface imagery shows a fully synthetic demonstration fixture, labeled as such in-product. Details covered by a pending invention disclosure are not shown here.
Evidence
- Working prototype, demonstrated and documented
- Nine interface captures from one running production-mode session against a fully synthetic fixture, each documented with the capability it shows and the limitation it must not obscure.
- How it is built
- A web interface (Next.js) over an analysis service (FastAPI, PostgreSQL with vector search), with a saved record of every analysis run, automated tests, and written decisions about where research data may and may not travel.
- Published method underneath
- The codebook pipeline implements the GATOS method: introduced in Humanities and Social Sciences Communications (2026), applied to 10,000+ posts in the Journal of Engineering Education (2025).
Related research
- How can researchers combine qualitative judgment with open-source generative AI to scale thematic analysis without hiding methodological choices?
- Thematic analysis with open-source generative AI and machine learning: A new method for inductive qualitative codebook development
- Using generative AI for large-scale qualitative analysis of social media posts to understand why people leave computer science
Maintenance
- Maintained by
- IDEEAS Lab
- Status
- Pilot and evaluation ahead
- Last reviewed
- August 28, 2026