There is lively debate over how far AI tools can or should be involved in research and evaluation. Yet AI is already embedded in both our methods and many of the interventions we evaluate, whether we like it or not. Reading the exchange between Jowsey and colleagues on one side, and two responses by Friese and colleagues on the other, I find myself nodding in agreement with parts of all three – and landing back on the “humans-first, AI-enhanced” spot I described earlier.
The open letter: drawing a line
In case you haven’t seen it, Jowsey et al. (2025) published an open letter, signed by hundreds of qualitative researchers from around the world, rejecting the use of generative AI in what they call “Big Q” qualitative approaches. They drew a red line around methods like reflexive thematic analysis, phenomenology, ethnography and similar reflexive approaches, arguing that GenAI should not be used in any phase of such work – not even coding.
Their reasoning rested on three big claims.
First, they argued that GenAI is a form of simulated intelligence: a statistical prediction engine without understanding of the world or the meaning of language, and therefore with no capacity to generate genuinely meaningful themes from qualitative data on its own. Anything it produces can at best resemble reflexive analysis, but never be it.
Second, they insisted that reflexive qualitative research is a distinctly human practice, done by humans with, about and for humans. The heart of this work is subjective, situated, psychodynamic meaning‑making, mindful of power and context. In their view, only humans can do that, so involving GenAI in the analytic process is methodologically incongruent.
Third, they pointed to the social and environmental harms entangled with GenAI: extractive data practices, energy and water use, e‑waste, land clearing, and the exploitation of low‑paid workers who moderate toxic content, documented in growing work on AI supply chains. For them, refusing GenAI is not just a methodological choice but an ethical obligation grounded in social and environmental justice.
I resonate with a lot of this. In my own writing, for example, I have stressed that large language models (LLMs) are trained to predict plausible‑sounding text, not to seek truth, and that they do not (yet) possess understanding, consciousness or intent in a human sense. When I ask an AI assistant to write about a topic I know well, I can instantly see if its first pass is not quite right – and more often than not, I have to correct it, which is precisely why I think we should be wary of trusting it uncritically in domains where we lack that same sniff test. Seen in that way, the letter’s insistence that GenAI is not, and should not be treated as, a reflexive analyst in its own right feels right to me, even if some techno-optimists might argue that GenAI does in effect ‘understand’ and ‘evaluate’ for many practical purposes.
Friese’s response: fighting the wrong enemy?
Friese’s initial single-author response (2025) took the open letter seriously, but pushed back on the idea that “GenAI cannot make meaning” automatically implies “therefore we must refuse it altogether”. She argued that the letter was, in effect, fighting a straw man, an imaginary world in which qualitative researchers simply hand over interpretive authority to a chatbot and call the result reflexive analysis.
Her central claim was that when researchers remain firmly in charge – with AI literacy and methodological awareness – GenAI can function as a thinking companion or scaffold rather than as a replacement analyst. She described AI as a way of externalising some of the cognitive work we already do with highlighters, memos, diagrams and peer conversations: surfacing patterns, suggesting contrasts, and helping us ask better questions of the data, while leaving interpretive authority squarely with the human researcher.
Friese’s new 2026 article on moving from coding to conversation further develops this thought, arguing that AI enables more dialogic forms of interpretation and knowledge creation, and that it doesn’t just speed analysis up but opens new ways of doing it.
Friese (2025) also pointed out that qualitative research has long drawn on ideas emphasising that meaning‑making is not a purely internal, isolated human process, even when humans remain the ones ultimately doing the interpreting. We already “think with” tools, texts, spaces and artefacts; in that light, she suggested, it is odd to single out GenAI as uniquely disqualifying, provided we keep it in the role of non‑agentic cognitive artefact rather than giving it epistemic authority.
On the ethical side, she agreed that environmental and labour harms are real, but questioned whether categorical abstinence by a relatively small group of researchers will change the trajectory of Big Tech. Instead, she argued for responsible use, regulatory advocacy, and collective efforts to push for more sustainable and just AI infrastructures – a harm‑reduction lens rather than a politics of refusal.
As someone who has been exploring “AI‑enhanced evaluation”, Friese’s framing is close to how I have been thinking about AI in my own work. In my earlier Substack, I argued that AI can be thought of as analogous to an e‑bike: it augments our effort, but it still needs a human rider who chooses the destination, the route and the level of assistance, who steers the bike and ultimately remains accountable. Friese is, in a sense, making the same argument for qualitative analysis: AI can help us pedal further, faster, and explore new routes to the destination, but it must not be left to ride off on its own - or things can go very wrong very quickly.
A second response: beyond binaries
A further open letter by Friese, Nguyen‑Trung, Powell and Morgen (2025) then joined the conversation. They responded directly to Jowsey et al., but with a different emphasis from Friese’s single‑author article.
Where the first response was a detailed critique of the original open letter’s arguments and use of literature, the four‑author letter focused on rejecting a simple “pro‑AI vs anti‑AI” binary. They argued that categorical refusal can marginalise efforts to design and use GenAI in reflexive, ethically informed ways, and that qualitative researchers should instead explore “researcher‑led and responsible use” under conditions of strong methodological literacy, transparency, governance and still-emerging good practice.
This second letter also argued that meaning in qualitative research is always created in interaction between people, their tools and their environments, and that GenAI can be treated as a powerful but essentially tool‑like aid to thinking, rather than as an independent interpreter. On this view, GenAI supports human analysis but does not replace the researcher’s responsibility for interpretation and ethics. That is close to my own “humans-first, AI‑enhanced” stance, though I am perhaps more cautious about how quickly practice and governance might catch up with the technology.
How this sits with my own work
So where do I land, as an Evaluation and Value for Investment practitioner and academic who also values qualitative and mixed-methods work and has been a curious experimenter with AI?
In my AI‑enhanced evaluation piece, I used the economic idea of comparative advantage to frame the division of labour between humans and AI. Even if AI ends up being “better” at a wide range of tasks in some abstract sense, there are still functions where humans’ distinctive strengths and responsibilities matter most, and others where AI can do the job at a much lower opportunity cost. The sweet spot is to have humans and AI specialise in the things each does best, and to coordinate that effort deliberately. If we get the mix right, then in theory we should achieve better, quicker evaluation that more people can afford, boosting demand for and use of evaluation rather than making human evaluators redundant.
I argued that AI has a comparative advantage in assisting (under human responsibility) with things like:
Rapid scoping, drafting and redrafting, and as a brainstorming assistant
Rapid research and synthesis (if we stay sceptical and double‑check sources)
Large‑scale pattern detection across diverse datasets
Automating routine tasks such as transcription and data cleaning.
Areas where I see humans having a comparative advantage include:
Curiosity, including the drive to explore ambiguity, challenge assumptions and ask better questions
Relationships and trust‑building with stakeholders, including the subtle social skills required for good interviews, workshops and deliberation
Contextual understanding – attuning to cultural, political and historical nuance, reading body language and tone, noticing the significance of conflicting worlds of evidence and outliers
Creativity and adaptive thinking in complex, unpredictable environments
Ethics and accountability, including reflecting on power, equity and potential harms, and being answerable to codes of conduct and communities
Sense‑making and storytelling – probing the “why” and “so what” behind patterns, and co‑constructing meaning with the people affected
Evaluative reasoning – using explicit criteria, standards and evidence to make value judgements that are not just technically tidy, but are arrived at collectively and deliberatively, and are socially and politically responsible.
Those human strengths line up quite well with what Jowsey et al. described as the “meaning‑making core” of reflexive qualitative research. I share their view that this core should not be automated, though we don’t have to completely exclude AI involvement - quite to the contrary, I expect AI may enhance more and more sub-tasks within these domains as the tools continue to improve, while we humans stay in the driver’s seat.
When I wrote about AI being able to bash out a decent descriptive report on monitoring data or satisfaction surveys, I argued that this is not evaluation, because it lacks explicit human values and value judgements.
Instant coffee isn’t a bad drink. Its only crime is the use of the name “coffee”. It should have a different name to avoid misleading. Similarly, “evaluator-free evaluation” may at times provide useful and valid analysis - it just shouldn’t use the name “evaluation”.
The same argument applies here: a GenAI‑produced thematic analysis that has not been interrogated and directed through human reflexivity should not, in my view, be labelled reflexive qualitative research, because mis-labelling risks confusing users and gradually crowding out the very practices it imitates, however technically sophisticated the output becomes.
Where I part company with the Jowsey et al. open letter is their conclusion that, since GenAI should not be our meaning-maker and is entangled with harm, we should keep it away from reflexive qualitative work altogether. That doesn’t sit comfortably with my experience or with the principles I have been sharing in this Substack series. To me it seems practically unworkable: even if we did want to refuse AI entirely, the funders, implementers, and data environments many of us work in are already heavily AI‑mediated. Evaluators operate within those systems, and the choice is between promoting ethical and effective use or leaving it to others.
In my Gell‑Mann amnesia metaphor, I worried about the way we can oscillate between seeing AI’s flaws clearly on topics we know well, and then forgetting those flaws when we work in less-familiar domains. The antidote I suggested there was not to ban AI, but to cultivate disciplined scepticism and to design our workflows in ways that reduce automation bias and anchoring. For example:
Think first, then ask AI – so the machine’s draft doesn’t bias our thinking before we have articulated our own view
Use AI to sense-check our thinking, making human-AI “sparring” a two-way street for detecting faulty logic, unclear arguments, missing nuances, factual errors, etc.
Treat AI outputs as rough drafts, hypotheses, or simulations to be tested, not as answers to be accepted
Keep humans in charge for everything that involves values, ethics, relationships, context, power, and real impacts on people’s lives.
Those are the kinds of guardrails that, I think, make space for something like what Friese is proposing: GenAI as a conversation partner, pattern amplifier and external sketchpad for abductive reasoning, sometimes better than us at particular tasks, albeit with variable reliability, but still not a black‑box oracle or evaluator.
My position, for now
Putting all this together, here is where I currently stand.
I agree with Jowsey and colleagues that GenAI does not and cannot “understand” qualitative data in the same way humans do, and that it shouldn’t, by itself, perform reflexive qualitative analysis. For me, the key issue is not whether future systems might approximate human-level analysis (increasingly, they can), but that they would still not occupy the social roles of researcher, evaluator or rights-bearing agent who can be held to account.
For clarity, I am trying to keep three questions separate. First, there is capability: in some domains, GenAI already performs at or above typical human levels, and that trend will likely continue. Second, there are epistemic questions about what, if anything, it means to say that such systems “understand” or “interpret” qualitative data. Third, there are social and ethical roles: who is recognised as a researcher or evaluator, who can enter into relationships with participants, and who can be held accountable for harms. My own humans-first stance is mostly about this third category, even as the first two keep evolving.
I also agree with Jowsey and colleagues that qualitative work, especially in justice‑oriented fields, needs to take seriously the environmental and labour harms tied up with AI. As evaluators, we should be among the people interrogating those impacts.
I agree that AI introduces new risks of harm to participants, misinterpretation of marginalised voices, flawed value judgements, and numerous other bads - which humans are already pretty good at too, tbh.
The real danger, I think, is not AI as such, but naive expectations and use – treating GenAI as an interpretive subject rather than as a fallible tool. A blanket ban risks shutting down thoughtful innovation and pushing AI use underground, rather than fostering the AI literacy and reflexive practice we urgently need.
Ultimately, if my hunch about comparative advantage is right, so that AI-enhanced, human-led evaluation can get us more, better, quicker, cheaper, then at what point does it become unethical not to use AI tools in evaluation?
Reading these pieces together, I see three broad positions in play: a politics of refusal that draws a hard red line around reflexive qualitative work (Jowsey et al.), a detailed methodological critique of that refusal (Friese’s solo response), and an invitation to move beyond binary “for or against AI” framings towards researcher‑led, reflexive use under tight conditions (Friese et al.).
My perspective sits closest to that third space, along with a strong emphasis on keeping meaning‑making, evaluative judgement and accountability as non‑delegable human responsibilities. In my own work, I want to keep exploring AI‑enhanced reflexive evaluation and qualitative analysis under some fairly strict conditions:
I want humans – evaluators, researchers, stakeholders – to do the first round of thinking, framing and sense‑making, so GenAI is responding to our questions rather than generating the agenda.
I want AI to be used, transparently, for what it does well: accelerating background tasks, suggesting patterns and questions, stress‑testing our frameworks, and helping us see our blind spots – always with humans checking, challenging and revising.
I want evaluative judgements and meaning‑making about people’s lives to remain non‑delegable human responsibilities, grounded in relationships, context and values, even if AI helps us organise and interrogate the material.
I want the profession to invest in AI literacy, ethics and governance, so we can make informed choices about which tools to use, how to minimise harm, and when to say “no, this task belongs with people”.
Human-factors research in aviation has a similar lesson: when pilots mainly monitor highly reliable automated systems, vigilance tends to drop and automation-induced complacency and error can creep in; when systems are designed so that automation also monitors human performance and keeps pilots actively engaged, overall safety improves. That is roughly how I am thinking about AI in evaluation: we should not sit back and watch the machine work, but design workflows where AI and humans explicitly check and challenge each other.
In other words, I do not want AI to be our evaluator or our reflexive qualitative analyst. I want it to be a powerful, slightly unreliable, sometimes brilliant, occasionally annoying research assistant – an e‑bike for our minds – that lets us go further and faster on the bits it is good at, while we stay firmly at the controls when it comes to meaning, values, equity and judgement.
The AI field is evolving rapidly, and I believe holding views lightly is an important feature of evaluative thinking. For both of these reasons, my views about AI in evaluation may change as AI tools, regulatory environments, and evidence about benefits and harms evolve, including potential shifts in our collective acceptance of what counts as good-enough AI support for different kinds of work.
Thanks for reading!
And thanks to Steve Powell for helpful peer review. Opinions, errors, and omissions are my responsibility.
References
Jowsey, T., Braun, V., Clarke, V., Lupton, D., & Fine, M. (2025). We reject the use of generative artificial intelligence for reflexive qualitative research. Qualitative Inquiry. https://doi.org/10.1177/10778004251401851
Friese, S. (2025). Response to: “We reject the use of generative artificial intelligence for reflexive qualitative research.” SSRN. https://doi.org/10.2139/ssrn.5690262
Friese, S., Nguyen‑Trung, K., Powell, S., & Morgen, D. (2025). Beyond binary positions: Making space for critical and reflexive GenAI integration in qualitative research. SSRN. https://dx.doi.org/10.2139/ssrn.5962174
Friese, S. (2026). From coding to conversation: A new methodological framework for AI-assisted qualitative analysis. Violence Against Women. Advance online publication. https://doi.org/10.1177/10778004251412871



Postscript: Given the topic, it feels important to be transparent about how I used AI in this article.
First, I read the Jowsey et al. letter and the two responses myself and jotted down my own reactions – where they align, diverge, or build on each other and on my earlier writing about AI‑enhanced evaluation.
I then gave Perplexity Pro my initial notes plus the three articles (and some of my prior blog posts), and asked it to suggest an outline. After refining that outline, I prompted it to draft the text one section at a time.
From there, I went through and picked over every word: editing, deleting, and rewriting so that the tone, arguments and judgements were squarely mine and that I was representing the authors’ positions fairly and with nuance.
Once I had a full draft in my own voice, I asked Perplexity to review it for the overall arc, logical coherence, balance between the three pieces, and clarity of argument. I then did another round of revisions and sent it to a human peer reviewer (thanks, Steve!) before fussing over it some more and pressing “publish”.
In other words, I used AI as an assistant for outlining, drafting and stress‑testing – not as an evaluator, analyst or arbiter of meaning. As with all my Substack posts, regardless of what tools I use, you can hold me accountable for every word.
I appreciate how other people are making visible their small & big decisions about using AI in their work. This is an example from the commercial research sector https://report.stripepartners.com/2026-gen-ai-research-playbook/index.html?utm_source=linkedin&utm_medium=thought_leader_post&utm_campaign=2026_genai_x_research_playbook&utm_content=Tom_Hoy