AI as the SPSS, R, and Stata for Qualitative Research

Quantitative researchers spent forty years learning to delegate computation to software without surrendering judgment. Qualitative researchers now face the same transition, with a tool that answers in the register of analysis rather than in numbers. A delegation framework, a project structure, and a reading of where the methodological dispute currently stands.
AI
qualitative research
workflow
methodology
reproducibility
Author

Kostadin Kostadinov

Published

August 20, 2026

Modified

August 25, 2026

When quantitative researchers opened SPSS for the first time in the 1980s, nobody asked whether the software was doing statistics “for real.” The question was whether the researcher understood what the software was doing: whether the output made sense, whether the model was appropriate, whether the assumptions held. The computer handled the arithmetic and the human handled the judgment, and that division of labour became so uncontroversial that it eventually became invisible.

Qualitative research is now entering an equivalent transition, with one important difference: the division of labour is much harder to draw. A large language model can code transcripts, propose themes, surface patterns across a corpus, draft summaries and identify contradictions, and it does all of this in prose that resembles analysis rather than in a column of coefficients. That resemblance is exactly why qualitative researchers need to be more deliberate about delegation, not less. Once you can no longer tell whether you are operating a tool or handing over a task, you have already lost track of who is doing the thinking.

The frame I find most useful is the one that worked for statistical software. An LLM is an instrument: it extends your capacity to process text in the way that R extends your capacity to process numbers. No competent quantitative researcher lets R choose the regression model, and no competent qualitative researcher should let an LLM choose the interpretation. The whole skill lies in knowing where that boundary sits, and — as I will argue below — a substantial part of the field currently denies that the boundary can be drawn at all.

NoteThe short version

AI in qualitative research sits roughly where SPSS sat in 1990: a powerful instrument that most users have not yet learned to wield with discipline. Structure the project so that human judgment and machine computation never occupy the same step, delegate the mechanical work, retain the interpretive work, and document what was delegated to whom. Whether any delegation at all is compatible with reflexive methodologies is an open dispute, and taking a position in it is now part of the methods section.

1. The analogy, and where it strains

The comparison to statistical software is more than a rhetorical device; the structural parallel is close enough to be worked systematically.

Quantitative tools Qualitative AI tools
What they handle Arithmetic, matrix algebra, optimisation, randomisation Text processing, pattern matching, pattern suggestion, summarisation
What they cannot do Choose the model, interpret the estimate, judge causal structure Interpret lived experience, understand cultural context, determine what something means
What they changed Removed manual computation; freed time for design and interpretation Removes manual coding labour; frees time for interpretation and reflexivity
The main risk Treating software output as findings rather than as intermediate results Treating AI-generated codes as findings rather than as intermediate results
The core skill Knowing which model fits your question Knowing which tasks can be delegated without changing what the methodology claims to know

R does not understand epidemiology, and GPT-5 does not understand phenomenology. In both cases the researcher’s task is to know the method well enough to recognise when the tool is doing the wrong thing, and fluent output does not relax that requirement.

The analogy strains at two points, and both matter. The first is register. A coefficient and a standard error are obviously intermediate quantities; nobody mistakes them for a finding. A theme statement written in competent academic English invites exactly that mistake, and the invitation is psychological rather than technical, which is why the safeguards below are procedural. The second is that the arithmetic SPSS performs is deterministic and inspectable, whereas an LLM’s mapping from transcript to code is stochastic, opaque and unstable across runs. A regression can be recomputed by hand; a coding decision produced at temperature 0.7 cannot be reconstructed even by the model that produced it. Reproducibility in the SPSS sense is not available here, and pretending otherwise is the most common overclaim in this literature.

2. What the evidence actually shows

The empirical base has grown quickly and unevenly. The most informative synthesis to date is the scoping review by Kempny and colleagues, published in BMC Medical Research Methodology in 2026, which screened 4,201 de-duplicated records from five databases covering January 2020 to May 2025 and included 75 peer-reviewed empirical studies that used an LLM at a substantive stage of qualitative analysis.

Three findings from that review are worth carrying into practice. The first concerns concentration: OpenAI GPT models appeared in 93% of the included studies, so the accumulated evidence describes the behaviour of one model family rather than of LLMs in general, and it is not obvious how far it transfers. The second concerns where the tools are actually being applied. Coding assistance (43 studies) and theme identification (41 studies) dominate, and thematic analysis is by a wide margin the host methodology (38 studies), followed by content analysis (12). In other words, the most frequent application is also the one where the methodological objections are sharpest, a mismatch I return to in the next section.

The third finding is the one I would put in front of anyone who cites headline accuracy figures. Reported agreement between LLM and human coders ranged from 36% to 99%. A range that wide is not a performance estimate; it is an indication that performance is governed almost entirely by moderators — task complexity, prompt construction, the rigour of the validation procedure — that the primary studies mostly failed to document. The reporting deficit is severe on the review’s own accounting: 75% of studies reported no parameter settings at all, only 13 reported temperature, 12 reported context length, 4 reported top-p, and 45% did not even specify whether the model was accessed through an API, a web interface or a local deployment. Against that, 95% discussed ethical considerations and 97% claimed human verification of AI output — figures high enough to suggest that verification is being asserted rather than described. When almost every paper says a human checked the output and almost no paper says what checking consisted of, the claim has stopped carrying information.

What survives this is not a performance benchmark but a gradient, and the gradient is architectural rather than incidental. Performance tracks task explicitness: the more structured, descriptive and rule-governed the task, the better these models do, and the more contextual, interpretive and culturally situated the task, the worse. This follows directly from what the models are — next-token predictors trained on a corpus that over-represents some registers, languages and populations and under-represents others — and it is therefore not the kind of limitation that the next model generation dissolves.

3. The disagreement worth taking seriously

It would be convenient to present careful delegation as the settled consensus position. It is not, and a methods section written as though it were will not survive an informed reviewer.

In December 2025, Jowsey, Braun, Clarke, Lupton and Fine published a position statement in Qualitative Inquiry, co-signed by 419 qualitative researchers from 32 countries, rejecting the use of generative AI for reflexive qualitative research outright. Their argument runs on three tracks. Methodologically, they hold that reflexive approaches such as reflexive thematic analysis constitutively require a subjective, positioned, reflexive researcher, so that a statistical text generator cannot perform the analysis rather than merely performing it badly. Epistemically, they argue that reflexive qualitative work is a human practice undertaken by humans, with or about humans, and that the models tend to reproduce dominant framings while flattening marginal ones. Ethically and environmentally, they point to resource extraction, emissions, and the labour conditions of the largely Global South workforce that trains content filters — costs that are external to the analysis but not external to the researcher’s responsibility.

The first two tracks deserve engagement rather than dismissal. If a methodology defines the researcher’s situated subjectivity as the analytic instrument, then delegating interpretation is not a shortcut but a category error, and the tool’s accuracy is beside the point. The third track is the one most easily set aside by researchers who would not set aside an equivalent argument about, say, fieldwork travel, and it is the one I find hardest to answer honestly.

The response has been quick and not uniform. Greenhalgh, writing in the same journal in 2026, declined to sign, on the grounds that a binary of adopters against refusers obscures the question that matters. Her distinction is between tools that displace, obscure or constrain the researcher’s reflexive engagement with the data and tools that do not; qualitative research has always used non-reflexive instruments — transcription software, search functions, coding frameworks — without anyone claiming that the software understood meaning. She reframes the ethical objection as a governance problem, concerning the scale of use, institutional accountability and who bears the environmental cost, and argues that governance questions are better addressed by staying in the conversation than by withdrawing from it. De Paoli has argued a stronger version of the same objection to categorical rejection.

My own position is closer to Greenhalgh’s, and I hold it with less confidence than the volume of the debate might suggest. What the dispute does establish is that the delegation decision is no longer a private workflow preference. It is a methodological commitment that has to be stated and defended, and the defence has to be made in the vocabulary of the methodology rather than in the vocabulary of efficiency.

4. Methodological congruence comes first

The practical consequence is that the sequence matters. The question is never “which tasks can AI do well?” but “which tasks can this methodology delegate without changing what it claims to know?” — and the second question is answered before any tool is opened.

The three families of thematic analysis illustrate how much latitude varies within what is nominally one method. Coding reliability approaches, which treat codes as a measurement instrument and inter-coder agreement as evidence of quality, are the most permissive: a second coder is by design interchangeable, and an LLM applying a fixed codebook is a plausible second coder whose agreement can be quantified in the ordinary way. Codebook approaches occupy a middle position, in which the codebook is a human theoretical product but its application is largely classificatory. Reflexive thematic analysis is the most restrictive, because the researcher’s subjectivity is the analytic resource rather than a source of error to be controlled, and there is no coherent sense in which a model can be a second reflexive analyst. Phenomenological, narrative and participatory designs restrict delegation further still.

Prahl’s AI-Reflexivity Checklist offers a usable procedure for making this determination before rather than after the fact. It asks five questions of the planned task — how far it is purely descriptive, how far meanings vary across contexts, how much experiential nuance is at stake, what the ethical exposure is, and whether the output is fully reversible — and sorts the task into one of three human-in-the-loop postures: delegate, assist, or human-led. Reversibility functions as a hard constraint rather than one criterion among five, on the sensible ground that an irreversible output cannot be audited or repaired. The checklist takes perhaps twenty minutes and produces a short record that a reviewer can read, which is more than most of the 75 studies in the Kempny review managed.

5. The project structure

What follows is a practical layout for an AI-assisted qualitative study. It follows the logic of numbered R scripts in a quantitative project: everything has a place, every step is traceable, and the sequence can be re-run to check that the same result emerges. The one structural departure is that AI interactions live in dedicated logs rather than inside the analysis scripts, so that the machine’s contribution remains separable after the fact.

2026-patient-experience-study/
├── README.md                          # study overview, team, status
├── protocol/
│   ├── study_design.md                # RQ → epistemology → methodology
│   └── ai_use_plan.md                 # pre-declared: what AI will and will not do
├── data/
│   ├── raw/                           # original transcripts, never edited
│   ├── deidentified/                  # cleaned for AI processing
│   └── processed/                     # analysis-ready extracts
├── scripts/
│   ├── 01_transcription.R             # or notes on transcription workflow
│   ├── 02_data_prep.R                 # cleaning, de-identification, formatting
│   ├── 03_human_familiarisation.md    # researcher notes from initial reading
│   ├── 04_coding.R                    # your coding framework (human-led)
│   └── 05_themes.R                    # theme development and refinement
├── ai/
│   ├── prompts/                       # every prompt you used, versioned
│   │   ├── deductive_coding_prompt.md
│   │   ├── theme_suggestions_prompt.md
│   │   └── negative_case_prompt.md
│   ├── logs/                          # raw AI outputs, timestamped
│   │   ├── 2026-07-15_coding_batch1.md
│   │   └── 2026-07-16_theme_review.md
│   └── audit_trail.md                 # model, version, date, settings, decisions
├── analysis/
│   ├── codes/
│   │   ├── codebook.md                # final codebook with definitions
│   │   └── code_comparison.csv        # human vs AI code assignments
│   ├── themes/
│   │   ├── theme_map.svg
│   │   └── theme_definitions.md
│   └── quotations/
│       └── verified_quotes.csv        # participant, location, original text, code, decision
├── output/
│   ├── figures/
│   └── tables/
├── manuscript/
│   └── manuscript.qmd
└── logs/
    ├── reflexivity_log.md             # evolving researcher reflections
    └── ai_reflexivity_log.md          # how AI shaped what you noticed

The two files most likely to be skipped are the two that carry the methodological weight. The AI use plan is written before the tool is opened; it records the congruence judgment from the previous section, states which tasks will be delegated, and gives the reason. The AI reflexivity log is maintained during analysis and records how the model’s suggestions moved the interpretation — where they were accepted, where they were rejected, and on what grounds. Without both, there is no audit trail, and the distinction between AI-assisted and AI-generated becomes impossible to demonstrate to anyone who asks.

6. The delegation framework

The decision rule maps onto the SPSS analogy. Delegate the work you would hand to a research assistant with clear instructions; keep the work that requires your expertise as the analyst; treat everything in between as requiring documented supervision.

Mechanical tasks

Transcription refinement is data preparation rather than analysis: cleaning the output of automatic speech recognition, fixing diarisation, standardising punctuation. So is formatting and organisation, whether that means converting interview notes into a standard extract format, normalising file names, or building structured data files out of messy raw input. Applying a pre-existing, human-developed codebook to excerpts is classification, and the qualitative equivalent of running table(); it is also the best-supported application in the empirical literature, though the wide agreement range in §2 should temper expectations about how well it will work on any particular corpus. Corpus searching — locating every instance where a topic, word or concept appears across a large dataset — is grep with better recall on paraphrase, and nothing more. Negative case searching exploits the same retrieval strength: the model is good at finding excerpts that sit awkwardly with a proposed theme, and you are the one who decides whether the friction is meaningful. Comparing two coding schemes for overlap and divergence is bookkeeping.

Interpretive tasks

Reading the transcripts is not delegable, and the reason is epistemological rather than conventional: interpretation without prolonged engagement is not interpretation. An AI summary of a transcript you have not read gives you the model’s compression of the data in place of your familiarity with it, and you will not be able to tell the difference from the inside.

Developing the initial coding framework is a human intellectual product whether it is derived deductively from theory or inductively from the data. A model can propose candidate codes; the framework, and the theoretical commitments it encodes, are yours. The same holds for the interpretive judgments that constitute the analysis. When a participant says “It was fine, I suppose,” deciding whether that registers resignation, politeness or satisfaction is the analysis. A model will give you an answer, and the answer will be a plausible guess drawn from the distribution of similar sentences in its training data, rather than a reading grounded in your relationship with the participant, your knowledge of the setting and your sense of what the pause before “I suppose” was doing.

Illustrative quotations are never copied from an AI output. They are retrieved from the original transcript against a participant and line identifier, because these models alter wording, silently repair grammar, merge excerpts from different speakers, and produce fluent quotations that do not exist in the source. Final theme development and theorisation are the intellectual core of the study rather than a text-processing problem. Reflexive accounting — noting how your position, history and relationship to the data shape what you noticed — is irreducibly first-person.

The grey zone

Inductive coding sits in between: a model can propose an initial set of codes, but every proposal has to be checked against the transcript, which makes it a suggestion engine rather than a coder. Codebook development is similar, in that pattern highlighting is useful and the resulting structure is still a human product. Translation produces a serviceable first draft and a poor final one, because idiom, register and the distance between what someone said and what they meant require bilingual human judgment. Synthesis across many interviews is the most seductive case: models identify surface commonalities well, and the move from what people said to what it means is precisely the analytic work that cannot be handed over.

Task Posture Condition
Transcription refinement, formatting, corpus search Delegate Routine verification
Deductive coding against a human codebook Delegate Agreement quantified against independent human coding
Negative case search, cross-codebook comparison Delegate Researcher judges relevance of what is retrieved
Inductive coding, codebook development Assist Every proposal checked against the transcript
Translation, cross-interview synthesis Assist Bilingual or analytic human verification
Familiarisation, interpretation, quotation selection Human-led No delegation
Theme development, theorisation, reflexivity Human-led No delegation

7. A worked example

Suppose the study concerns how general practitioners in Bulgarian small towns experience the transition to electronic health records. Twenty-four semi-structured interviews have been conducted and transcribed in Bulgarian, each running 45 to 70 minutes.

Step 1: Establish methodological congruence

Before any tool is opened, four things go into protocol/ai_use_plan.md. The research question: how do GPs in Bulgarian small towns experience the transition to electronic health records, and what does it mean for their clinical practice? The epistemology: constructionist, since the interest is in how participants construct meaning rather than in counting occurrences of predefined categories. The methodology: reflexive thematic analysis. And the implication, which is the point of the exercise — reflexive TA treats the researcher’s subjectivity as a resource rather than a bias, so mechanical support is admissible and the interpretive work that constitutes the method is not. A codebook-driven design would permit considerably more delegation; this one does not. Given the Jowsey position statement, a study of this design also has to state whether it accepts any AI involvement at all, and defend the answer.

Step 2: Transcription and preparation

Audio is transcribed with Whisper or an equivalent, then reviewed against the recording, with particular attention to Bulgarian medical terminology, regional dialect features, and the hedging and irony that automatic transcription reliably flattens. Transcripts are then de-identified — names, practice locations, patient identifiers — with originals kept in data/raw/ and cleaned versions in data/deidentified/. Transcription is delegated; verification is not.

Step 3: Human familiarisation

Every transcript is read in full, with analytic notes taken on patterns, surprises and contradictions, and written into scripts/03_human_familiarisation.md. For 24 interviews this is typically 20 to 30 hours of work, and it is the foundation for everything that follows.

WarningDo not skip this step

Uploading 24 transcripts and asking for themes is the qualitative equivalent of running every possible regression and reporting the ones that reached significance. The requirement for prolonged engagement is not a ritual; interpretation without familiarity is fabrication with a methods section attached.

Step 4: AI-assisted first-pass coding

Only now is the model engaged, and on a subsample. For each of five interview excerpts, a prompt supplies the de-identified text and the codebook, and asks the model to apply the codebook, identify the relevant passage for each code, and state its reasoning:

“Apply the following codebook to these interview excerpts. For each code, identify the relevant passage and explain your reasoning.”

You supply the codebook rather than asking the model to generate one. This is the deductive coding case, and the one with the strongest empirical support.

The model’s assignments are then compared against your own independent coding of the same excerpts, and the comparison is recorded in analysis/codes/code_comparison.csv:

Participant Excerpt Code AI assigned Human assigned Agreement
GP-07 “I still write on paper when the system is slow” Workaround Yes Yes
GP-12 “The system is fine, I suppose” Resignation No Yes

Disagreements drive refinement of the code definitions, the model is re-run on a fresh batch, and the loop continues until the framework stabilises. This is inter-rater calibration in all but name, and it is defensible because the model functions as a second coder under human final authority rather than as the coder. Record the agreement rate and where the disagreements concentrated; that number is the substance behind the phrase “human verification,” and the Kempny review suggests that almost nobody currently reports it.

Step 5: Theme development

With the codebook stable and all 24 interviews coded under human final authority, theme development proceeds iteratively. This is the interpretive core, and the model’s role here is adversarial rather than productive: ask it to propose alternative organisations of the codes, or to identify the excerpts most likely to embarrass an emerging theme.

“Here are 12 codes from my analysis. Suggest three alternative ways they could be organised into themes. For each, identify which codes cluster naturally and which seem forced.”

Then discard the suggestions you find unconvincing and sit with the ones that surprise you. The value lies in the friction against your own emerging reading rather than in the model’s answer, and the friction is worth logging even when the suggestion is rejected.

Step 6: Quotation verification

Every quotation in the manuscript is retrieved from the original transcript rather than from an AI output, and analysis/quotations/verified_quotes.csv records the retrieval:

Participant Transcript file Line/page Original text (verbatim) Code Theme Included in manuscript
GP-07 GP-07_transcript_bg.txt L142–144 “Когато системата падне, минавам на хартия — така сме свикнали” Workaround Adapting under pressure Yes

Fabricated, truncated, merged and quietly improved quotations are the best-documented failure mode of these systems, and an unverified quote in a published paper is a correction waiting to be issued.

Step 7: Audit trail and reflexivity

ai/audit_trail.md records the model and access route (say, GPT-5 via the API), the date range, the sampling parameters — temperature, maximum tokens, and the fact that transcripts were processed individually rather than batched — the exact system prompt, the versioned list of task prompts, and a summary of what was changed or rejected during coding reconciliation. On the Kempny figures, supplying this places a study in a small minority.

logs/ai_reflexivity_log.md records the interpretive traffic in both directions. Two entries from a study of this kind might read: “AI coding of GP-07 and GP-12 suggested I was over-coding ‘resignation’; reviewed the transcripts and found the excerpts were better described as pragmatic acceptance, and revised the code definition.” And: “The model’s alternative theme organisation grouped ‘training’ with ‘support’. I disagree, because the data show these are experienced differently, and I have noted the disagreement in the theme development log.” The point of the log is that it captures the cases where the machine changed your mind, which are the cases a reader has no other way of seeing.

8. The reporting standard

The COREQ+LLM extension, whose protocol was published in JMIR Research Protocols in September 2025 and which is registered as a guideline under development with the EQUATOR Network, will eventually supply an agreed checklist. Until it does, the working standard is that a methods section should let another researcher reconstruct what the researchers did, what the model did, what information the model received, what it produced, how the humans evaluated that output, and what ultimately determined the interpretation. If a reader cannot distinguish your analysis from one in which the model did the thinking, the reporting is not yet adequate.

A minimal disclosure has six components, and they are short. Name the tool and version, with the access route and date — “GPT-5 via the OpenAI API, accessed July 2026” — together with the sampling parameters, since these are omitted in three quarters of published studies and determine the reproducibility of nothing else in the paper. State the task: first-pass deductive coding of interview excerpts against a human-developed codebook. State the input: de-identified Bulgarian-language excerpts of 300 to 500 words, supplied individually. State the output: proposed code assignments with a rationale for each excerpt. State the human evaluation in quantitative terms, including who reviewed the output, who retained final authority, what the disagreement rate was, and where the disagreements clustered — “18%, concentrated in codes requiring cultural or emotional interpretation” carries information in a way that “all output was reviewed by the research team” does not. Finally, state the reflexive position: that use was pre-declared in an analysis plan, that a reflexivity log was maintained, and how the model’s suggestions bore on theme development.

The graduation of reporting effort follows the delegation posture. Mechanical delegation — transcription, formatting, deductive coding, corpus search — is reported directly and briefly. Grey-zone delegation — inductive coding suggestions, candidate themes, translation drafts — is reported with the qualification that human evaluation was applied to every output, and with the evidence for it. Any interpretive task delegated to a model is reported with an explicit methodological justification, and that justification has to answer the position statement in §3 rather than ignore it.

9. The multilingual dimension

Almost all of the evidence discussed above comes from English-language studies using GPT-family models, and the Kempny review names the limited evidence for non-English contexts as one of the field’s open problems. For Bulgarian-language research, the extrapolation is unsafe in a specific way.

Bulgarian is a morphologically rich Slavic language with a complex verbal aspect system, an analytic nominal system unusual among the Slavic languages, an extensive clitic inventory, evidential verb forms that mark whether the speaker witnessed what they report, substantial dialectal variation, and conventions of hedging, indirectness and politeness that carry considerable interactional weight. The evidential system alone should give pause: the distinction between what a GP saw happen and what they were told happened is grammatically marked in Bulgarian and disappears entirely in English translation. A model that reaches 80% agreement with human coders on English interviews has not thereby been shown to do anything in particular on Bulgarian data, and very few people have checked.

The safe workflow for multilingual work runs from original-language human interpretation, through AI assistance, to bilingual verification. The unsafe one runs from AI translation to English-only analysis, and its defect is that the interpretive step becomes invisible by being relocated into preprocessing. You believe you are analysing what participants said; you are analysing what the model rendered them as having said, and the two diverge exactly where the analysis is most interesting. Where translation is unavoidable, pilot the model on a small subset and compare its coding of the Bulgarian original against its coding of the translated text; the size of the gap is a study finding in its own right, and worth reporting.

10. Common failures

Failure Why it happens How to prevent it
Automation anchoring The model’s first-pass coding structure becomes the researcher’s frame Read the transcripts before engaging the model; keep analytic notes from familiarisation
Quotation fabrication Plausible-sounding quotes are generated that do not exist in the data Retrieve every quote from the original transcript against a participant and line identifier
Theatrical validation “AI output was reviewed by researchers”, with no detail Specify what was compared, by whom, and with what disagreement rate
Epistemological mismatch Delegating interpretation in reflexive TA or phenomenology, where subjectivity is the method Settle congruence before selecting tasks; document the judgment
Missing audit trail “ChatGPT was used to assist analysis”, with nothing further Maintain the use plan, audit trail and reflexivity log throughout
Parameter silence Model, version, temperature and deployment route go unreported Record them at the time of use, not at the time of writing
Privacy breach Identifiable transcripts uploaded to consumer services De-identify before processing; use enterprise or local deployment for sensitive data
Language assumption English-language performance assumed to transfer Pilot on a subset; compare coding of original against translated excerpts

11. Where this is going

Statistical software did not replace statisticians. It changed what statisticians spend their time on: the arithmetic was delegated, and design, interpretation and judgment became more consequential rather than less. The researchers who did well out of the transition were those who understood the method deeply enough to notice when the software was doing the wrong thing.

Something similar seems likely here, with two qualifications. The transition is faster, and the instrument is far less transparent than a statistical package. An R output displays the coefficient, its standard error and the degrees of freedom, and every step from data to estimate can in principle be reconstructed. A model output displays a theme statement that reads like something a thoughtful colleague might have written, and you have to determine for yourself whether it is one, with no access to how it was produced. That asymmetry makes methodological literacy more important rather than less, and it makes the apparatus described here — the separation of human and machine tasks, the audit trail, the reflexivity log — part of the method rather than administrative overhead sitting on top of it.

The second qualification is that a substantial body of qualitative researchers holds that the analogy fails at its root: that the arithmetic SPSS took over was never constitutive of statistical reasoning, whereas interpretation is constitutive of qualitative analysis, so that delegating it changes the object rather than the labour. I do not think that argument settles the matter, since it applies with full force to interpretive delegation and with very little force to transcription cleanup. But it does establish that the burden of justification sits with the person delegating, and it should be discharged in the methods section rather than in a footnote.

NoteThe question to answer

“Can AI do qualitative analysis?” is the wrong question. The useful one is which tasks can be delegated without changing what this methodology claims knowledge to be — a methodological question rather than a technical one, and the one that separates rigorous AI-assisted research from AI-generated prose with the surface features of analysis.


References

De Paoli S. Why we should reject to reject the use of generative artificial intelligence in qualitative analysis: a response to Jowsey, Braun, Clarke, Lupton, and Fine (2025). Qual Inq. 2026. doi:10.1177/10778004261425137

Fehring L, Frings J, Rust P, Kempny C, Thürmann PA, Meister S. Extension of the Consolidated Criteria for Reporting Qualitative Research Guideline to Large Language Models (COREQ+LLM): protocol for a multiphase study. JMIR Res Protoc. 2025;14:e78682. doi:10.2196/78682

Greenhalgh T. Reflexive qualitative research and generative AI: a call to go beyond the binary. Qual Inq. 2026. doi:10.1177/10778004261429383

Jowsey T, Braun V, Clarke V, Lupton D, Fine M. We reject the use of generative artificial intelligence for reflexive qualitative research. Qual Inq. 2025. doi:10.1177/10778004251401851

Kempny C, Frings J, Rust P, Meister S, Fehring L. The use and methodological reporting of large language models in qualitative research: a scoping review. BMC Med Res Methodol. 2026;26(1). doi:10.1186/s12874-026-02913-1

Prahl A. The AI-Reflexivity Checklist (ARC): a pre-analysis pause for LLM-assisted coding. Qual Health Res. 2026;36(2-3):181–190. doi:10.1177/10497323251401503

ImportantColophon

Written in August 2026. Empirical claims about model performance and reporting practice are drawn from the scoping review cited above; the project structure and delegation framework reflect my own experience integrating these tools into epidemiological and health services research, and should be read as a proposal rather than as a validated instrument. Corrections and disagreement are welcome at kostadinr.kostadinov@mu-plovdiv.bg.