IDEEAS Lab

CATS QDA

A local-first workspace where qualitative researchers analyze sensitive interview data with AI assistance they can inspect, override, and audit.

Built on our work in AI for qualitative research: Using Large Language Models and Generative AI to Scale Qualitative Data Analysis.

Research prototype, internal, pilot and evaluation ahead.

The problem

Rigorous qualitative coding, memoing, and theme development take substantial researcher time, so valuable interview and document collections are often analyzed only at limited scale, and the usual AI shortcuts hide exactly the judgments researchers need to see.

An end-to-end analysis workspace that runs entirely on the researcher’s own computer: transcripts, codebooks, themes, memos, search, and data-grounded chat, with agentic workflows that propose codes and organize themes while recording an execution trace the researcher can review stage by stage.

For qualitative researchers, mixed-methods researchers, education and health researchers.

What it looks like

CATS QDA project workspace showing a synthetic study with research questions, data file, code, theme, and participant counts.
A synthetic qualitative-analysis workspace with research questions and live project counts. The counts come from a deliberately small documentation fixture, not a study result.
Turn-level transcript review with speaker roles, dialogue acts, filtering, and version controls.
Turn-level transcript review with confirmed speaker roles, dialogue acts, filtering, and version controls. The transcript is fictional text; this does not demonstrate live transcription quality.
A data-grounded chat response citing source turns from the synthetic interview corpus, with tool usage reported.
A data-grounded synthesis citing source turns from the synthetic corpus. The displayed response is a seeded successful run, not a claim of general model reliability.
A five-stage codebook generation execution trace from corpus understanding through quality review.
An execution trace from corpus understanding through candidate generation, novelty analysis, and theme organization. The run was seeded as a completed example; it does not establish autonomous validity.

Research basis

The workbench builds the GATOS codebook-generation method, published in Humanities and Social Sciences Communications and applied at scale in the Journal of Engineering Education, into an interactive system. The pipeline a researcher watches run (corpus understanding, candidate generation, novelty analysis, theme organization, quality review) is the published method, made inspectable.

What it does today

  • Organizes projects around research questions, with transcript ingestion, speaker attribution, dialogue acts, filtering, and version controls.
  • Answers researcher questions from the corpus with grounded chat that cites source turns and reports which retrieval tools ran.
  • Generates candidate codebooks through a multi-stage agent pipeline and keeps a persistent execution trace of every stage.
  • Structures codebooks with definitions, abstraction levels, theme grouping, and explicit acceptance state: the researcher accepts or rejects, not the model.
  • Links analytic memos to files, codes, and themes with provenance labels.
  • Runs local-first, with the database, vector search, and language models on the researcher’s own machine, so sensitive corpora never leave it.

What it does not do

  • Not hosted and not installable without a developer-oriented local stack; explicitly single-user and loopback-only, with no authentication or multi-user isolation.
  • No completed external pilot or user study.
  • The quality indicators it reports (saturation, coverage, coherence, balance) have not been validated against benchmarks.

What it does not show yet

  • A working demonstration against synthetic data is evidence that the system exists and runs, not that it improves research quality, productivity, or coding validity. Those claims wait on the planned pilot and evaluation.
  • Agent-proposed codes and themes are candidates for researcher judgment, not findings. The design thesis is inspectability, and the honest corollary is that the human review it enables is still required.

Outcomes

  • Turns the lab’s published GATOS method into an interactive, inspectable platform.
  • Demonstrates full-stack research engineering: durable agent execution, provenance records, and grounded retrieval over a private corpus.

Limitations and responsible use

  • Use with real research data stays inside the documented local trust boundary and applicable IRB and data-use agreements; the platform makes analysis local, not consent unnecessary.
  • All interface imagery shows a fully synthetic demonstration fixture, labeled as such in-product. Details covered by a pending invention disclosure are not shown here.

Evidence

Working prototype, demonstrated and documented
Nine interface captures from one running production-mode session against a fully synthetic fixture, each documented with the capability it shows and the limitation it must not obscure. As of August 28, 2026.
How it is built
A web interface (Next.js) over an analysis service (FastAPI, PostgreSQL with vector search), with a saved record of every analysis run, automated tests, and written decisions about where research data may and may not travel. As of August 28, 2026.
Published method underneath
The codebook pipeline implements the GATOS method: introduced in Humanities and Social Sciences Communications (2026), applied to 10,000+ posts in the Journal of Engineering Education (2025). As of 2026.

Maintenance

Maintained by
IDEEAS Lab
Status
Pilot and evaluation ahead
Last reviewed
August 28, 2026

All tools