Two hundred sixty documents, an agent with no instructions, and the consultant who ran the real engagement to judge the result
An engagement billed at tens of thousands of euros. Two consultants, a few weeks, two hundred sixty documents. An AI rebuilt the final deliverable in a few hours, for the price of a Claude subscription. The document it produced was good, and almost nobody reading it alone would have challenged it. But it was by comparing that result with the real deliverable, alongside the consultant who actually ran the engagement, that we put our finger on what our clients really buy.
Two hundred sixty documents, no instructions
July 2026. Alenia had just billed an engagement worth a few tens of thousands of euros: a post-mortem after a series of incidents at a client, run over several weeks by two of my fellow consultants. I did not take part in that engagement. But I collected everything it contained: the commercial proposal, the documents the client provided, the emails, the transcripts of the interviews conducted in the field. About two hundred sixty documents. And I asked Opus 4.8, at high effort, with Cowork agents running in parallel, to redo on its own what two consultants had spent weeks producing.
The protocol was deliberately minimal: access to the full folder, no instructions on form or substance, and a single directive: figure it out.
Two clarifications matter here, because they change what this experiment actually demonstrates. First, these two hundred sixty documents were not a raw pile: my colleagues had already sorted through everything the client had sent and selected the relevant pieces. Second, the documents included the transcripts of the interviews my colleagues had conducted. So the AI did not start from nothing. It benefited from filtering and collection work already done by humans: questions asked, follow-ups, and context gathered live. It writes, then, from field work that was filtered and collected, and that it did not do itself.
What the AI produced, minute by minute
Ten minutes for a first context document, built from several Cowork agents reading the two hundred sixty documents in parallel. That step consumed half of a five-hour quota. Eight more minutes for the detailed deliverable in text format: a fifth of the quota. Another quarter of an hour for a PowerPoint version, this time requested by me from the text document: more than 30% of the next quota, exhausted before the end, forcing me to wait for it to renew.
The first attempt at formatting in the Alenia template failed. I had given it the wrong template, the one for a commercial proposal. Twenty minutes and another 28% of quota for an unusable result. The second attempt, with the right template this time, took 4% of quota.
In total: about twenty-five minutes of my time, spread over five hours of waiting between quota windows. I hardly retouched anything by hand. In the moment, it impressed me more than it reassured me. All of it for the price of a monthly subscription. The original engagement had cost tens of thousands of euros and several weeks of work for two people.

And what does the person who ran the real engagement say?
I showed the deliverables to one of the two consultants who had actually run the engagement. He could judge; I could not. I had not read the 260 original documents. The AI could have told me anything plausible without my noticing.
His verdict: average to good. The incident timeline, the root cause analysis, the five successive whys followed a structure very close to what the real engagement had produced. The context document, written with no instructions at all, would have been enough to help someone discovering everything from scratch understand the subject.
The PowerPoint, on the other hand, remained too abstract. It did not anchor its claims and did not emphasize what really mattered. On one point in the folder, no document covered a key step of the engagement: the AI filled the gap with sweeping strides rather than flagging the absence of a source. And it got the importance of a technical system mentioned at the opening of the folder wrong, one detail among two hundred sixty documents, not the crux of the problem. It could not know that yet.
I thought the real obstacle would be there, in these errors of judgment about what matters. But I was wrong.
The question that shows no mercy
When my colleague presented the real post-mortem, he found himself facing the client’s executive and the head of the area concerned. He knew, live, who had said what, when, and at what level of the hierarchy, because he had conducted the interviews himself. An executive at that level systematically probes the point that clashes with what they already believe. You then have to answer on the source, on its reliability, on the context in which it was gathered. Immediately.
Faced with the same document produced by the AI, my colleague could not have answered those questions. Not because the document was wrong. Because he had not written it himself, had not asked the questions, had not heard the hesitations and contradictions as they were being expressed.
No bluffing is possible.
Transcripts give the words. They do not give the memory of the scene in which those words were spoken.
What can’t be delegated?
The structuring difference is therefore here. Formatting and writing can be delegated without much loss: that is already what my colleagues did on the real engagement, before this experiment, to refine a set of questions, to retrieve a piece of information from Teams exchanges, to write a clean paragraph in English from raw ideas. Collecting and analyzing information, on the other hand, cannot be delegated without loss. Information you have not collected and challenged yourself cannot be recovered after the fact, even transcribed in black and white. Diving back in to make it your own would demand an effort that is nearly impossible to carry out seriously.

Without this minimum of collection, challenge, and analysis, nobody can present or defend a document with authority. You become disconnected from the ground, and in front of a sponsor who pushes, it shows immediately: you are no longer defending a diagnosis, you are bluffing. Bypassing this work remains technically possible, as the experiment just showed. It is still counterproductive, and increasingly so, because presenting and defending the document is becoming decisive for convincing a client.
And the real engagement confirms it: the written document was not what convinced the client. What counted was the conviction deployed orally, on a few strong points. The document has value. It is never enough on its own, without someone to defend it.
What this changes for Alenia in how we use AI
The answer is not to replace the steps of an engagement with AI, but to keep all of them and speed them up with AI: refining an interview guide before conducting the interviews, synthesizing faster after conducting them, formatting a deliverable once the consultants themselves have done the analysis. Skipping a step remains counterproductive, because you lose contact with what you produce, and the ability to defend it.
This conclusion is somewhat uncomfortable. It assumes that a consulting engagement will keep looking, in its structure, like what it is today, only faster at each step. That may not be the right answer in the long run.
Another avenue deserves raising, without claiming to settle it here. What if the deliverable were not doomed to remain a fixed presentation, which a consultant brings and defends alone in front of the sponsor? It could become a cleaned-up, structured body of documentation that the sponsor queries live, with the AI retrieving on demand the source and justification for every contested claim. The consultant’s role would then shift: no longer defending every line in person in front of the client, but steering the AI during collection and analysis, then ensuring that what it tells the sponsor live stays anchored in a real source rather than in a gap filled with sweeping strides, like the one observed on this folder. This model moves the problem more than it solves it: it assumes we can trust the AI to say honestly that it does not know, instead of pressing ahead with confidence on information that does not exist. Nothing in this experiment yet proves that it already knows how to do that.
For now, my one certainty is what cannot be delegated to AI. Collecting the information remains the cognitive price to pay in order to defend what you produce. And that is what our clients buy without always knowing it: not the document, but someone able to hold their ground in front of them the day they challenge it.