Subscribe to GEN
Login to GEN
Add a Comment
Every mail vendor is now selling the same promise: an inbox that reads itself, sorts itself, summarises itself and, if you will just tick one more box, replies to itself. The demonstrations are genuinely impressive. Somebody sends a sprawling twelve-message thread, the assistant reduces it to five bullets, drafts a courteous response, and the audience makes the noise the vendor paid for.
What the demonstration never shows is the moment three weeks later when a customer produces one of those replies in a dispute and asks the obvious question: did you mean this, or not?
There is a real, usable, substantial productivity gain sitting in the mail client. We use it ourselves, every day. But it lives entirely on one side of a line, and the line is not a matter of taste. It is the line between a tool that assists a person who then makes a decision, and a system that takes over and makes the decision instead.
An email sent from your mailbox is your statement. There is no footnote in English contract law for "the assistant wrote that one".
Both are marketed as "AI email". They are not the same product, and they do not carry the same risk:
The gap between them is one button and roughly the whole of your legal exposure. Almost everything valuable about AI in email sits in the first category. Almost everything that has ever gone badly wrong sits in the second.
People treat email as informal because it feels informal. The courts do not share that view, and they have not shared it for a long time. An exchange of emails can form a binding contract, vary an existing one, waive a right, extend or blow a deadline, and satisfy a "notice in writing" clause. None of that requires anybody to have intended it to be a formal document.
The point is made rather sharply by Neocleous v Rees [2019] EWHC 2462 (Ch), in which the High Court held that an automatically generated email sign-off, the footer the client inserts on its own without the sender doing anything at all, was capable of amounting to a signature for the purposes of section 2 of the Law of Property (Miscellaneous Provisions) Act 1989. A block of text nobody chose to type, appended by software, was enough to help bind a party to a disposition of land.
Now consider what a generative model appends to an email, and how much more than a signature block it is prepared to say.
The other well-known illustration comes from Canada rather than England, but it is instructive because of the defence that was run. In Moffatt v Air Canada (2024 BCCRT 149), an airline chatbot invented a bereavement fare policy that did not exist. The airline argued, in writing, that the chatbot was a separate legal entity responsible for its own actions. The tribunal was unimpressed, and the airline paid. We wrote about that case and the DPD incident at the time in 20240521, and nothing since has improved the position.
So, plainly:
And this is before we reach the ordinary commercial reality that if you promise something in an email, most businesses will simply expect you to honour it rather than litigate about it. The cheaper failure is the one where you quietly absorb the cost of a promise you never made.
Most take-over failures never get near a court. They are merely mortifying, and they are far more frequent:
That last one deserves emphasis. Payment diversion fraud works by inserting itself into a real conversation at a plausible moment. An automated responder has no scepticism, no sense of occasion, and no memory of the phone call you had with that supplier in March. It is a machine for being agreeable, pointed directly at the people whose job it is to be suspicious.
Everything above concerns what the machine says. There is a second problem, and in regulatory terms it is the larger of the two, because it applies even when the assistant never sends a single message and never gets anything wrong.
A corporate mailbox is one of the densest concentrations of personal data any business holds. Names, job titles, direct dial numbers, home addresses, bank details, sickness absence, grievances, references, medical notes forwarded by an employee who did not think, immigration status, salary discussions, and the entire correspondence history of every customer and supplier you have. Threads run for years and collect participants nobody remembers adding.
Under Article 4(2) of the UK GDPR, processing includes "disclosure by transmission". Sending the body of an email to a third party so that its model can generate a summary is processing of personal data by that third party, on your instruction. It is not a neutral technical operation, and calling it a feature does not change its character.
So the questions arrive in a specific order, and they are uncomfortable:
Your customer emailed your sales address. They did not agree that six years of your correspondence with them would be transmitted to a company in Virginia so that it could be turned into a vector.
If the provider is outside the United Kingdom, this is a restricted transfer under Chapter V, and it needs adequacy regulations or an approved mechanism such as the International Data Transfer Agreement or the UK Addendum to the standard contractual clauses, supported by a transfer risk assessment and whatever supplementary measures that assessment identifies. This is the same discipline any competent organisation already applies to its sub-processors. An AI feature does not get a pass on it because it arrived as part of a licence bundle.
For the two jurisdictions that host almost every model on the market, the assessment is not a formality:
It is also worth remembering that the shape of the regime shifts. The Data (Use and Access) Act 2025 has adjusted parts of the UK framework, principally around automated decision-making and research, but it did not relax the international transfer rules and it did not reduce the accountability obligations. Anything you rely on today, you should expect to have to justify again.
There is a further exposure that has nothing to do with data protection law at all, and it catches people who have their GDPR position perfectly in order.
Read the confidentiality clause of almost any non-disclosure agreement, framework agreement or professional engagement letter you have signed. It will restrict disclosure of the other party's confidential information to third parties without prior written consent, usually with a narrow carve-out for professional advisers and staff on a need-to-know basis. There is no carve-out for language models.
When a member of staff asks an assistant to summarise the negotiation thread on a deal governed by that agreement, the contents of that thread have been disclosed to a third party. Nobody intended a breach, nobody will notice, and the position is not improved by the fact that the disclosure was automated. If your organisation handles privileged material, tender responses, patient or client information, or anything under a duty of confidence, this is the clause that bites first, and it bites regardless of jurisdiction.
A breach of confidence is the survivable version of this problem. It is a claim, and claims settle. What follows does not.
The first thing to understand is that almost everything turns on which tier of a provider's service the text reached, and that the tier is very often chosen by whichever member of staff installed something, rather than by anyone accountable:
Microsoft 365 Copilot deserves its own paragraph, because it is the one most likely to be already switched on in a business reading this. It runs OpenAI models on Microsoft's own infrastructure rather than sending your tenant's data to OpenAI, and Microsoft states that prompts, responses and the Graph data behind them are not used to train foundation models. Take that at face value. Then notice what is not being claimed: Copilot interactions are retained, stored against the user's mailbox and surfaced through Purview, audit logging and eDiscovery. Your administrators can see them, which means so can an opposing party's solicitors on a disclosure order. Every prompt an employee ever typed, including the ones phrased with total candour about a customer, a colleague or a claim, is now a discoverable business record that did not exist before.
The Chinese providers deserve a more careful answer than they usually get, because two quite different things are routinely muddled together. Sending your mail to a Chinese hosted API is subject to everything above and rather less recourse: no adequacy finding, no enforceable commitment a UK business could realistically rely on, and a statutory duty to assist state intelligence work sitting behind the provider. The honest position is not that we know what they do with the context, it is that the assessment cannot be completed, and processing you cannot assess is processing you cannot authorise.
The weights are an entirely different proposition, and we will come back to them, because some of the best models you can run on your own hardware were trained by Chinese laboratories and given away.
"We do not train on your data" and "we do not keep your data" are different sentences. Notice which one you have been given, and for which tier.
Even the good tier, with the good contract, rests on an assumption that ought to trouble anyone who has read a disclosure order.
During the litigation between the New York Times and OpenAI, a US court ordered OpenAI to preserve output log data that its own published policy said would be deleted, including content covered by its standard retention commitments. The order was contested, and it was later narrowed. That is not the point. The point is that a promise a British business had relied on was suspended by a foreign court, in proceedings that business was not party to, could not influence, and in most cases did not know were happening until it read about them.
This is what a transfer risk assessment is actually for, and it is why the honest answer to "is it safe?" is never a link to a vendor's trust page. Your data protection position is contingent on a commercial commitment that a third party can be compelled to break, in a legal system where you have no standing.
Everything so far assumes the vendor behaves exactly as documented. In August 2025 two of them demonstrated, within three weeks of each other, that the more ordinary failure needs no bad faith at all. It only needs a product decision made by somebody who has never met you.
OpenAI had added a sharing feature to ChatGPT. You picked a conversation, generated a link, and, if you also ticked a box marked "Make this chat discoverable", search engines were permitted to index it. It was opt-in, it was documented, and it was catastrophically misunderstood, because people ticking a box in a share dialogue reasonably assumed they were sharing with the colleague they were sending the link to. Fast Company found thousands of those conversations sitting in Google's results, complete with names, email addresses, CVs, children's names and a great deal of material people had plainly typed in the belief that they were alone.
On 1 August 2025 OpenAI withdrew it. Dane Stuckey, the company's chief information security officer, described it as a short-lived experiment to help people discover useful conversations, and said this:
"Ultimately we think this feature introduced too many opportunities for folks to accidentally share things they didn't intend to, so we're removing the option. We're also working to remove indexed content from the relevant search engines."
That is a fair and honest response, and it is worth reading twice for what it concedes. The feature worked as designed. The terms were not breached. Nothing was hacked. The company simply shipped an experiment, users misjudged what a checkbox meant, and private conversations became public documents. "Working to remove indexed content" is also doing a great deal of quiet work in that sentence, because a page that has been in Google's index has been fetched, cached, screenshotted and scraped by parties who will not be removing anything.
Later the same month it happened again, and worse. Grok's share button published conversations to xAI's own website, where they were indexed by Google, Bing and DuckDuckGo. Forbes reported more than 370,000 conversations exposed. This time there was no discoverability toggle to misread and no warning that sharing meant publishing; users pressing a button labelled "share" to send a link to one person had published to the open web. Among the material recovered were medical and psychological questions, business details and at least one password.
Not a breach. Not a training corpus. Not a subpoena. A button, relabelled by somebody in another country, on a product your staff were already using.
Now apply that to a mail assistant. The content in question is not what one employee chose to type into a chatbot. It is the correspondence of every customer, supplier, candidate and colleague you have, with the reasonable expectation of confidence that attaches to a letter. If a feature can turn a chat into a search result, it can do the same to a summarised thread, and you will discover it the way everybody else did, which is by reading about it.
Here is the part that turns a compliance problem into an existential one.
Where content is used for training, it does not sit in a database from which it can be retrieved and removed. It is absorbed into the weights of a model, in a form nobody can index, audit or reverse. There is no deletion request that works, because there is nothing discrete to delete. The model has simply learned from it, and every subsequent model trained on that lineage may carry the same learning forward.
Nor is this only theoretical. Researchers have repeatedly demonstrated training data extraction against production models, recovering verbatim passages from training corpora, including personal data, using nothing more sophisticated than carefully chosen prompts. The models are not supposed to do this. They do it anyway, because memorisation is a side effect of how they are built rather than a defect somebody forgot to fix.
So consider what that means with a mailbox as the input. Your pricing model. The rates you agreed with one customer and not another. The counsel's advice on a dispute. The redundancy list before it was announced. The acquisition nobody has been told about. The clinical detail in a referral. The identity of the person who raised a whistleblowing concern.
A confidentiality breach is a claim you can settle, an apology you can make and an insurance policy you can call on. Your own commercially sensitive material re-emerging eighteen months later inside an answer given to a competitor who typed a plausible question is not a claim. It is an outcome. You will not be told it happened, you cannot test whether it has happened, and you will find out, if you ever find out at all, from somebody else.
Which is why "read the terms carefully" is the wrong conclusion to draw from all of this. Terms are revised, tiers are reorganised, features are shipped as experiments, defaults are changed at a product review you are not invited to, and companies are acquired. The only control that does not depend on somebody else's continued good behaviour is not creating the artefact in the first place.
Everything above is a reason to be careful. This is the part that makes it tractable, and it is refreshingly simple.
If the model runs on hardware you own, in a building you control, on a network you operate, then there is no third party. There is no disclosure by transmission to anyone. There is no processor to contract with under Article 28, no sub-processor to add to a list, no restricted transfer, no transfer risk assessment, no adequacy question, no data bridge to monitor, no foreign intelligence statute to assess, no preservation order in a case you have never heard of, no retention at somebody else's end, no training corpus, and no confidentiality clause to worry about, because nothing was disclosed to anybody. The processing is your own processing, under your own lawful basis, on your own infrastructure, exactly as it is when you run a mail server or a database.
An entire compliance workstream, several assessments, an annual review and a genuine residual risk are replaced by a decision about where a piece of software runs. That is not a small saving, and it is available today, for nothing.
Which brings us back to the point deferred earlier, and it is the most useful thing in this article.
An open-weight model is a file. You download it, you put it on a machine you own, and you run it. There is no account, no API key, no telemetry, no usage tier, no terms of service governing what you do with your own inference, and no relationship with the laboratory that trained it. It is a large numerical artefact sitting on a disk, in the same category as a database engine or a compiler.
And the striking thing about the current state of the field is that several of the strongest open-weight families were trained by Chinese laboratories and given away: Qwen from Alibaba, DeepSeek, GLM from Zhipu, Kimi from Moonshot, alongside Mistral in France, Gemma from Google and a healthy independent scene. Many are released under genuinely permissive terms, commonly Apache 2.0 or MIT.
Read that against the transfer analysis above, because the inversion is worth dwelling on. The same laboratory whose hosted endpoint you could not lawfully send a customer's correspondence to has handed you, at no cost, the means never to need it. Once the file is on your hardware there is no transfer, no processor, no jurisdiction question and nothing for anybody to compel, because there is no continuing relationship to compel anything through. The PRC National Intelligence Law has no more purchase on a set of weights running in your server room than the CLOUD Act has on a copy of PostgreSQL.
The objection to a Chinese model was never the model. It was the destination. Remove the destination and the objection goes with it.
Nor does this require a research budget. Summarising a thread, tagging by rule, translating a paragraph and generating embeddings are not frontier tasks. Small models handle them competently on a decent workstation, mid-sized ones on a single GPU in a server, and embedding models in particular are cheap enough to be an afterthought. For a great deal of routine mail work the quality gap against a hosted frontier model is narrow, and it is narrowing.
Four things worth knowing, two of which are usually got backwards:
And one thing matters more than all four: running the model locally makes the processing private, it does not make the output right. A local model will invent a delivery date with precisely the same confidence as a hosted one. It fixes the data protection problem completely and the accountability problem not at all, which is why the human stays on the send button either way.
Assume you accept all of the above and want an assistant rather than an autopilot. The next question is whose software it runs inside, and that turns out to decide almost everything else.
With vendor-integrated AI in webmail or Outlook, the answers are fixed for you:
This is the ecosystem argument we have made before in 20260102, applied to the one system every business actually depends on. It is also the sovereignty argument from 20260115 and 20260303, and the hosting argument from 20240529: the more of your correspondence is processed somewhere you cannot point to on a map, the less of it is genuinely yours, and the harder it becomes to answer a data subject who asks a perfectly reasonable question about where their information went.
If the assistant is part of the mail service, you have not bought an assistant. You have extended the lease on the mail service.
Thunderbird is generally filed under "the free one", which does it a considerable disservice. Judged as an enterprise mail client rather than as a curiosity, it holds up rather well:
policies.json file allow account settings, restrictions and add-on control to be deployed across a fleet without touching each workstation.Betterbird is worth knowing about separately. It is a downstream build of Thunderbird, maintained by a former Thunderbird developer, tracking the same release trains but carrying fixes and features that have not made it upstream. For organisations that want the platform without waiting on an upstream release cycle for a defect that affects them today, it is a serious option, and it is the same profile format, so moving between the two is not a migration. It is also the open source argument demonstrated rather than asserted: because the source was public, somebody who disagreed with a decision could take it and fix it, and everybody else got the benefit. With a closed client, a defect that affects you is a support ticket and a hope.
Being honest about the trade: with a desktop client you own the deployment, the profile strategy, the backup and the update discipline. That work does not vanish, it moves to you. It is also precisely the work we do for clients every day, and it is a far smaller burden than the one you take on by handing your entire correspondence history to a platform whose commercial terms you do not set.
The strongest argument for Thunderbird is not any single feature. It is that the feature set is not a fixed menu decided in another country by people who have never seen your business.
In a vendor-integrated client you get the tools the vendor built, in the order the vendor built them, and you campaign politely on a feedback forum for the rest. In Thunderbird you install what you actually need:
None of that requires a licence upgrade, a tenant-wide change, or a conversation with an account manager. It requires a decision by you about what your people need.
One caution, because it is the same caution we gave in 20251128: an add-on is a supply chain. Review what it does, check that it is maintained, understand the permissions it asks for, pin the versions you have tested, and use the enterprise policy file to control what may be installed on managed machines. Freedom to bolt things on is only an advantage if somebody is deciding what gets bolted on. The saving grace is that these are overwhelmingly open source projects, so "review what it does" is a real instruction that somebody can actually carry out, rather than the polite fiction it becomes when the feature is a closed component of a platform you are obliged to accept whole.
Which brings us to our own contribution. GENA is an extension for Thunderbird and Betterbird that puts an AI assistant inside the mail client rather than beside it: summarise a thread, rewrite a draft, generate content from an inline instruction, digest an inbox, tag messages against your own rules, translate, strip the formatting that other clients insist on injecting, and search an entire mailbox by meaning rather than by keyword.
Four of its design decisions matter more than the feature list, and each of them is a direct answer to a problem set out above.
GENA does not hard-code a provider. You configure the endpoint, the headers, the request body and the path to the generated text in the response. That makes it work with Ollama or LM Studio on a workstation, vLLM on your own GPU, or a hosted service if that is what a particular job needs.
This is not a technical preference, it is five protections at once, and the first of them is the one that keeps a data protection officer in the room:
Keyword search fails constantly, and it fails for a boring reason: it matches the words in the message, and you are trying to remember the subject of the message. You are looking for the enquiry about bulk water supply. The email says "tanker deliveries" and never once uses the word water.
GENA embeds messages as vectors, indexes them in Qdrant, and searches by meaning. New mail is indexed as it arrives and deleted mail is pruned on the next scan. You ask a question in plain English, the relevant messages are retrieved, and the model answers with a supporting evidence table you can click straight through into the message itself.
The evidence table is the part that matters. An answer with no sources is a machine asking to be believed. An answer with the six messages it came from is a research assistant, and you remain the person who reads them and decides. And because the index is yours, the entire corporate memory of the mailbox does not have to be shipped into somebody else's search product to make it findable.
GENA does not send email. It is a deliberate design constraint, not a missing feature.
Instead, you write the email, and where you want a paragraph generated you write an instruction on a line beginning with %. Press the GENA button and each instruction is replaced, in place, with generated content. The surrounding text is untouched, multiple instructions are processed in order, and nothing whatsoever leaves the client until you read the result and press send yourself.
In practice it looks like this:
Dear Mr Blobby, Thank you for your email. %explain why using Microsoft Outlook for corporate email is a poor choice and what we recommend instead If you have any questions, please let me know. Regards Rich
You have decided who to write to, what the email should say, what argument it should make, and what tone it should take. The model has done the typing for the one paragraph you did not want to type for the ninth time this month. That is a real saving, repeated many times a day, and it costs you nothing in accountability, because the accountability boundary is exactly where it always was: at the send button, with a person on the other side of it.
Compare that with an automatic responder. It also saves the typing. It also removes the person who would have noticed that the thread was hijacked, that the deadline stated in it was wrong, that the customer was not asking what the model thought they were asking, and that the paragraph promised a delivery date nobody in the business had agreed to.
| Assist (semi-manual) | Take-over (automatic) | |
|---|---|---|
| Who composes | You, with the machine drafting on instruction | The machine, from the thread |
| Who reads before sending | A person, every time | Nobody |
| Who is accountable | The sender, knowingly | The sender, unknowingly |
| Hallucinated commitment | Caught before it leaves | Sent, and binding |
| Hijacked or fraudulent thread | Human scepticism applies | Answered politely |
| Time actually saved | Most of it | Slightly more, at considerable cost |
The efficiency difference between the two columns is small. The risk difference is the whole article.
Everything above is a vendor making claims about its own software, which is precisely the category of statement this article has spent several thousand words advising you not to accept. So do not accept it. Check.
GENA ships as a .xpi, which is a zip file. Rename it, unpack it, and what is inside is the code that runs: plain JavaScript, HTML and CSS, not minified, not bundled, not compiled, with the developer's comments still in place. There is no build step between what we wrote and what your machine executes, and no separate source release you have to take on faith as matching the binary, because the shipped artefact is the source.
That turns every claim in this section into something testable in a text editor rather than something to be believed:
Your own security team can review it before it goes anywhere near a mailbox. Your auditor can be handed the file. If it does something you dislike, you can change it, and having reviewed a version you can pin that version across the fleet through the same enterprise policy file that controls everything else. None of that requires our cooperation, our permission, or our continued existence as a company.
You should not have to trust something with the whole of your correspondence. You should be able to read it.
Now ask the same questions of an assistant built into a mail platform. What is transmitted, to which region, under what retention, and does that change next quarter? The best available answer is a documentation page describing intended behaviour, written by the party with the least interest in your reading it closely. That is a description of a promise. The code is a description of a mechanism, and only one of the two can be checked.
If you are introducing AI into mail, this is the short version of the policy, and it fits on one side of paper:
AI in the inbox is genuinely useful, and the useful part is not the part being marketed hardest. Summarising, searching, drafting and tagging save real time for real people. Sending on your behalf saves a few more seconds and hands a statistical text generator the authority to speak for your company in writing.
And remember whose data you are being invited to be generous with. A mailbox is not your information, it is thousands of other people's information, held by you because they had to send it somewhere. They did not consent to it being transmitted abroad for a model to process, they were never asked, and in most cases they could not sensibly have been asked. That makes the choice of where processing happens a governance decision at board level, not a preference buried in a settings panel.
Weigh it accordingly, because this is one of the few decisions in IT that cannot be reversed. A bad server can be replaced, a bad supplier can be exited, a bad contract can be renegotiated. Text that has been absorbed into somebody else's model weights stays there, and no amount of later diligence gets it back.
Keep the assistant on the assist side of the line. Run it in a client you control rather than in a service that controls you. Choose your own model, and favour a local one, so that confidentiality, jurisdiction, capability and cost all remain your decisions. Index your own mail so you can find it by meaning without surrendering it to somebody else's search index. Prefer open source at every layer, the client, the extension and where you can the model, so that "what does this actually do with my mail" is a question answered by reading rather than by asking. And keep a human on the send button, because that is not a limitation of the technology, it is the entire point.
GEN has been working with machine learning since 1999, long before "AI" was a marketing word. We run our own compute clusters, train and fine-tune our own models, and write our own post-processing and guardrail layers rather than wrapping someone else's API and hoping for the best. We provide commercial AI solutions to enterprise customers globally, with confidentiality and security as a core part of what makes our systems valuable. Client prompts and embeddings are never stored, indexed, shared, captured, or used for training. We believe large language models are most useful when they are applied carefully and appropriately to real-world problems, not treated as something to stuff into a workflow and hope for the best.
--- This content is not legal or financial advice & Solely the opinions of the author ---