Processing...

 AI in the Inbox: Assist, Not Autopilot

The Curious Codex

             8 Votes  
100% Human Generated
2026-09-01 Published, 2026-09-02 Updated
8160 Words, 41  Minute Read

The Author
GEN Blog

Richard (Senior Partner) LinkedIn

Richard has been with the firm since 1992 and was one of the founding partners

 

AI in the Inbox: Assist, Not Autopilot

AI in the inbox: an assistant that drafts, and a human who sends

Every mail vendor is now selling the same promise: an inbox that reads itself, sorts itself, summarises itself and, if you will just tick one more box, replies to itself. The demonstrations are genuinely impressive. Somebody sends a sprawling twelve-message thread, the assistant reduces it to five bullets, drafts a courteous response, and the audience makes the noise the vendor paid for.


What the demonstration never shows is the moment three weeks later when a customer produces one of those replies in a dispute and asks the obvious question: did you mean this, or not?


There is a real, usable, substantial productivity gain sitting in the mail client. We use it ourselves, every day. But it lives entirely on one side of a line, and the line is not a matter of taste. It is the line between a tool that assists a person who then makes a decision, and a system that takes over and makes the decision instead.


An email sent from your mailbox is your statement. There is no footnote in English contract law for "the assistant wrote that one".

Two Very Different Products Wearing the Same Badge

Both are marketed as "AI email". They are not the same product, and they do not carry the same risk:

  • Assist. The machine summarises, drafts, translates, rewrites, tags and searches. Nothing leaves the building. A person reads the output, corrects it, and presses send. Accountability never moves.
  • Take-over. The machine composes and sends, triages and files, or answers on your behalf while you are away. A statement is made in your name that no human being read before it was made.

The gap between them is one button and roughly the whole of your legal exposure. Almost everything valuable about AI in email sits in the first category. Almost everything that has ever gone badly wrong sits in the second.


Why the Distinction is Contractual, Not Cosmetic

People treat email as informal because it feels informal. The courts do not share that view, and they have not shared it for a long time. An exchange of emails can form a binding contract, vary an existing one, waive a right, extend or blow a deadline, and satisfy a "notice in writing" clause. None of that requires anybody to have intended it to be a formal document.


The point is made rather sharply by Neocleous v Rees [2019] EWHC 2462 (Ch), in which the High Court held that an automatically generated email sign-off, the footer the client inserts on its own without the sender doing anything at all, was capable of amounting to a signature for the purposes of section 2 of the Law of Property (Miscellaneous Provisions) Act 1989. A block of text nobody chose to type, appended by software, was enough to help bind a party to a disposition of land.


Now consider what a generative model appends to an email, and how much more than a signature block it is prepared to say.


The other well-known illustration comes from Canada rather than England, but it is instructive because of the defence that was run. In Moffatt v Air Canada (2024 BCCRT 149), an airline chatbot invented a bereavement fare policy that did not exist. The airline argued, in writing, that the chatbot was a separate legal entity responsible for its own actions. The tribunal was unimpressed, and the airline paid. We wrote about that case and the DPD incident at the time in 20240521, and nothing since has improved the position.


So, plainly:

  • There is no "the AI did it" defence. The system is your agent, operating on your instructions, from your domain. What it said, you said.
  • A confident invention is a representation. If an automated reply states a price, a lead time, a warranty position or a service level that is not real, you have made a misrepresentation, and section 2(1) of the Misrepresentation Act 1967 puts the burden of showing reasonable grounds on you.
  • Contractual machinery does not pause for automation. Notice provisions, response windows, acceptance of variations and waiver by conduct all operate on what was sent, not on who or what composed it.
  • Disclosure is disclosure. An automated reply that quotes the wrong thread, or attaches the wrong document, has disclosed personal or confidential data. The UK GDPR does not have a category for accidental automated candour, and where a solely automated process produces a decision with legal or similarly significant effects for an individual, Article 22 has views of its own.

And this is before we reach the ordinary commercial reality that if you promise something in an email, most businesses will simply expect you to honour it rather than litigate about it. The cheaper failure is the one where you quietly absorb the cost of a promise you never made.


The Tier Below the Lawyers

Most take-over failures never get near a court. They are merely mortifying, and they are far more frequent:

  • A summary that confidently reverses the meaning of the one paragraph in the thread that mattered.
  • A drafted reply that adopts a chirpy transatlantic register in the middle of a formal complaint.
  • An auto-reply that answers the wrong participant, in the open, with the internal position.
  • A triage rule that files the only genuinely urgent message of the week as a newsletter.
  • An assistant that replies helpfully inside a hijacked thread and confirms a set of fraudulent bank details, which is precisely the attack we set out in 20260808.

That last one deserves emphasis. Payment diversion fraud works by inserting itself into a real conversation at a plausible moment. An automated responder has no scepticism, no sense of occasion, and no memory of the phone call you had with that supplier in March. It is a machine for being agreeable, pointed directly at the people whose job it is to be suspicious.


The Part Nobody Costed: Your Mailbox is Full of Other People's Personal Data

Everything above concerns what the machine says. There is a second problem, and in regulatory terms it is the larger of the two, because it applies even when the assistant never sends a single message and never gets anything wrong.


A corporate mailbox is one of the densest concentrations of personal data any business holds. Names, job titles, direct dial numbers, home addresses, bank details, sickness absence, grievances, references, medical notes forwarded by an employee who did not think, immigration status, salary discussions, and the entire correspondence history of every customer and supplier you have. Threads run for years and collect participants nobody remembers adding.


Under Article 4(2) of the UK GDPR, processing includes "disclosure by transmission". Sending the body of an email to a third party so that its model can generate a summary is processing of personal data by that third party, on your instruction. It is not a neutral technical operation, and calling it a feature does not change its character.


So the questions arrive in a specific order, and they are uncomfortable:

  • Whose data is it? Overwhelmingly, not yours. It belongs to the customer who wrote in, the supplier's accounts clerk, the candidate who applied, the employee who raised a concern. They are the data subjects, and none of them are in the room when the AI feature is switched on.
  • What is the lawful basis? Consent is almost certainly not it, and could not realistically be obtained: you cannot ask every person who has ever emailed you whether they agree to their correspondence being embedded by a model in another jurisdiction. That leaves legitimate interests under Article 6(1)(f), which is available, but only after a balancing assessment that has been written down and can be produced. Very few businesses running an AI mail feature have done one.
  • Have you told anybody? Articles 13 and 14 require the categories of recipient to be disclosed. If your privacy notice does not mention that correspondence may be transmitted to a named AI provider, your privacy notice is now wrong, and so is your Article 30 record of processing activities.
  • Is there a processor contract? Article 28 requires a written contract with specified terms before a processor touches personal data on your behalf. A tick box in an application's settings screen is not that contract, and a consumer account's terms of service is very much not that contract.
  • Is there special category data in there? Yes, there is. There always is. Article 9 sets a higher bar, and a mail assistant does not read a thread selectively to avoid it.
  • Does anything come back? Retention at the provider's end, including retention for abuse monitoring, is retention. So is any use of the content for training, and the distinction between a tier that trains on inputs and one that does not is a contractual term you should be able to point at, not an assumption.
Your customer emailed your sales address. They did not agree that six years of your correspondence with them would be transmitted to a company in Virginia so that it could be turned into a vector.

And Then There is Where It Goes

If the provider is outside the United Kingdom, this is a restricted transfer under Chapter V, and it needs adequacy regulations or an approved mechanism such as the International Data Transfer Agreement or the UK Addendum to the standard contractual clauses, supported by a transfer risk assessment and whatever supplementary measures that assessment identifies. This is the same discipline any competent organisation already applies to its sub-processors. An AI feature does not get a pass on it because it arrived as part of a licence bundle.


For the two jurisdictions that host almost every model on the market, the assessment is not a formality:

  • The United States. The UK Extension to the EU-US Data Privacy Framework, the "data bridge", is a genuine and usable route, but it covers only those US organisations that have self-certified to the framework and opted into the UK extension. Confirm your provider is on the list rather than assuming. Beyond that, section 702 of FISA and the CLOUD Act mean the transfer risk assessment has real work to do, and the framework itself rests on US executive arrangements that have been challenged before and can be altered by a change of administration rather than a change of law.
  • China, meaning the hosted services. There are no adequacy regulations for China. Article 7 of the PRC National Intelligence Law 2017 obliges organisations to support and cooperate with state intelligence work. That is the assessment you would have to write and sign before sending a customer's correspondence to an endpoint in Hangzhou, and we do not think it can honestly be completed. Note carefully that this is an objection to the destination, not to the model, and the distinction turns out to matter enormously.

It is also worth remembering that the shape of the regime shifts. The Data (Use and Access) Act 2025 has adjusted parts of the UK framework, principally around automated decision-making and research, but it did not relax the international transfer rules and it did not reduce the accountability obligations. Anything you rely on today, you should expect to have to justify again.


The Contractual Trap Underneath the Statutory One

There is a further exposure that has nothing to do with data protection law at all, and it catches people who have their GDPR position perfectly in order.


Read the confidentiality clause of almost any non-disclosure agreement, framework agreement or professional engagement letter you have signed. It will restrict disclosure of the other party's confidential information to third parties without prior written consent, usually with a narrow carve-out for professional advisers and staff on a need-to-know basis. There is no carve-out for language models.


When a member of staff asks an assistant to summarise the negotiation thread on a deal governed by that agreement, the contents of that thread have been disclosed to a third party. Nobody intended a breach, nobody will notice, and the position is not improved by the fact that the disclosure was automated. If your organisation handles privileged material, tender responses, patient or client information, or anything under a duty of confidence, this is the clause that bites first, and it bites regardless of jurisdiction.


What Actually Happens to the Text You Send

A breach of confidence is the survivable version of this problem. It is a claim, and claims settle. What follows does not.


The first thing to understand is that almost everything turns on which tier of a provider's service the text reached, and that the tier is very often chosen by whichever member of staff installed something, rather than by anyone accountable:

  • Consumer tiers train on your content by default. ChatGPT's free and paid consumer plans use conversations to improve the models unless a user finds and changes the setting. Anthropic changed its consumer terms in 2025 so that Claude conversations are used for training unless the user opts out, with retention extended to five years for those who leave it enabled. Google's consumer Gemini terms provide for human review and multi-year retention of reviewed conversations. These are the products your people already have open in another tab.
  • Commercial tiers say the opposite, and the difference is enormous. The OpenAI API, ChatGPT Enterprise, Anthropic's commercial and API offerings, and Google's paid and Workspace products all state that customer content is not used to train foundation models. That is a real and meaningful commitment. It is also a contractual promise about future conduct, not a physical property of the system, and it protects only the traffic that actually went through that tier.
  • Not trained on is not the same as not kept. These are two separate statements, and vendors are careful to make only the one they can. API content is typically retained for a period for abuse monitoring before deletion, with zero retention available only on request and only for some endpoints.

Microsoft 365 Copilot deserves its own paragraph, because it is the one most likely to be already switched on in a business reading this. It runs OpenAI models on Microsoft's own infrastructure rather than sending your tenant's data to OpenAI, and Microsoft states that prompts, responses and the Graph data behind them are not used to train foundation models. Take that at face value. Then notice what is not being claimed: Copilot interactions are retained, stored against the user's mailbox and surfaced through Purview, audit logging and eDiscovery. Your administrators can see them, which means so can an opposing party's solicitors on a disclosure order. Every prompt an employee ever typed, including the ones phrased with total candour about a customer, a colleague or a claim, is now a discoverable business record that did not exist before.


The Chinese providers deserve a more careful answer than they usually get, because two quite different things are routinely muddled together. Sending your mail to a Chinese hosted API is subject to everything above and rather less recourse: no adequacy finding, no enforceable commitment a UK business could realistically rely on, and a statutory duty to assist state intelligence work sitting behind the provider. The honest position is not that we know what they do with the context, it is that the assessment cannot be completed, and processing you cannot assess is processing you cannot authorise.


The weights are an entirely different proposition, and we will come back to them, because some of the best models you can run on your own hardware were trained by Chinese laboratories and given away.


"We do not train on your data" and "we do not keep your data" are different sentences. Notice which one you have been given, and for which tier.

The Promise Is Only as Durable as the Vendor's Ability to Keep It

Even the good tier, with the good contract, rests on an assumption that ought to trouble anyone who has read a disclosure order.


During the litigation between the New York Times and OpenAI, a US court ordered OpenAI to preserve output log data that its own published policy said would be deleted, including content covered by its standard retention commitments. The order was contested, and it was later narrowed. That is not the point. The point is that a promise a British business had relied on was suspended by a foreign court, in proceedings that business was not party to, could not influence, and in most cases did not know were happening until it read about them.


This is what a transfer risk assessment is actually for, and it is why the honest answer to "is it safe?" is never a link to a vendor's trust page. Your data protection position is contingent on a commercial commitment that a third party can be compelled to break, in a legal system where you have no standing.


And the Controls Are Theirs to Change

Everything so far assumes the vendor behaves exactly as documented. In August 2025 two of them demonstrated, within three weeks of each other, that the more ordinary failure needs no bad faith at all. It only needs a product decision made by somebody who has never met you.


OpenAI had added a sharing feature to ChatGPT. You picked a conversation, generated a link, and, if you also ticked a box marked "Make this chat discoverable", search engines were permitted to index it. It was opt-in, it was documented, and it was catastrophically misunderstood, because people ticking a box in a share dialogue reasonably assumed they were sharing with the colleague they were sending the link to. Fast Company found thousands of those conversations sitting in Google's results, complete with names, email addresses, CVs, children's names and a great deal of material people had plainly typed in the belief that they were alone.


On 1 August 2025 OpenAI withdrew it. Dane Stuckey, the company's chief information security officer, described it as a short-lived experiment to help people discover useful conversations, and said this:

"Ultimately we think this feature introduced too many opportunities for folks to accidentally share things they didn't intend to, so we're removing the option. We're also working to remove indexed content from the relevant search engines."

That is a fair and honest response, and it is worth reading twice for what it concedes. The feature worked as designed. The terms were not breached. Nothing was hacked. The company simply shipped an experiment, users misjudged what a checkbox meant, and private conversations became public documents. "Working to remove indexed content" is also doing a great deal of quiet work in that sentence, because a page that has been in Google's index has been fetched, cached, screenshotted and scraped by parties who will not be removing anything.


Later the same month it happened again, and worse. Grok's share button published conversations to xAI's own website, where they were indexed by Google, Bing and DuckDuckGo. Forbes reported more than 370,000 conversations exposed. This time there was no discoverability toggle to misread and no warning that sharing meant publishing; users pressing a button labelled "share" to send a link to one person had published to the open web. Among the material recovered were medical and psychological questions, business details and at least one password.


Not a breach. Not a training corpus. Not a subpoena. A button, relabelled by somebody in another country, on a product your staff were already using.

Now apply that to a mail assistant. The content in question is not what one employee chose to type into a chatbot. It is the correspondence of every customer, supplier, candidate and colleague you have, with the reasonable expectation of confidence that attaches to a letter. If a feature can turn a chat into a search result, it can do the same to a summarised thread, and you will discover it the way everybody else did, which is by reading about it.


And If It Does Reach a Training Corpus, There Is No Delete

Here is the part that turns a compliance problem into an existential one.


Where content is used for training, it does not sit in a database from which it can be retrieved and removed. It is absorbed into the weights of a model, in a form nobody can index, audit or reverse. There is no deletion request that works, because there is nothing discrete to delete. The model has simply learned from it, and every subsequent model trained on that lineage may carry the same learning forward.


Nor is this only theoretical. Researchers have repeatedly demonstrated training data extraction against production models, recovering verbatim passages from training corpora, including personal data, using nothing more sophisticated than carefully chosen prompts. The models are not supposed to do this. They do it anyway, because memorisation is a side effect of how they are built rather than a defect somebody forgot to fix.


So consider what that means with a mailbox as the input. Your pricing model. The rates you agreed with one customer and not another. The counsel's advice on a dispute. The redundancy list before it was announced. The acquisition nobody has been told about. The clinical detail in a referral. The identity of the person who raised a whistleblowing concern.


A confidentiality breach is a claim you can settle, an apology you can make and an insurance policy you can call on. Your own commercially sensitive material re-emerging eighteen months later inside an answer given to a competitor who typed a plausible question is not a claim. It is an outcome. You will not be told it happened, you cannot test whether it has happened, and you will find out, if you ever find out at all, from somebody else.


Which is why "read the terms carefully" is the wrong conclusion to draw from all of this. Terms are revised, tiers are reorganised, features are shipped as experiments, defaults are changed at a product review you are not invited to, and companies are acquired. The only control that does not depend on somebody else's continued good behaviour is not creating the artefact in the first place.


Why a Local Model Collapses the Entire Problem

Everything above is a reason to be careful. This is the part that makes it tractable, and it is refreshingly simple.


If the model runs on hardware you own, in a building you control, on a network you operate, then there is no third party. There is no disclosure by transmission to anyone. There is no processor to contract with under Article 28, no sub-processor to add to a list, no restricted transfer, no transfer risk assessment, no adequacy question, no data bridge to monitor, no foreign intelligence statute to assess, no preservation order in a case you have never heard of, no retention at somebody else's end, no training corpus, and no confidentiality clause to worry about, because nothing was disclosed to anybody. The processing is your own processing, under your own lawful basis, on your own infrastructure, exactly as it is when you run a mail server or a database.


An entire compliance workstream, several assessments, an annual review and a genuine residual risk are replaced by a decision about where a piece of software runs. That is not a small saving, and it is available today, for nothing.


Which brings us back to the point deferred earlier, and it is the most useful thing in this article.


An open-weight model is a file. You download it, you put it on a machine you own, and you run it. There is no account, no API key, no telemetry, no usage tier, no terms of service governing what you do with your own inference, and no relationship with the laboratory that trained it. It is a large numerical artefact sitting on a disk, in the same category as a database engine or a compiler.


And the striking thing about the current state of the field is that several of the strongest open-weight families were trained by Chinese laboratories and given away: Qwen from Alibaba, DeepSeek, GLM from Zhipu, Kimi from Moonshot, alongside Mistral in France, Gemma from Google and a healthy independent scene. Many are released under genuinely permissive terms, commonly Apache 2.0 or MIT.


Read that against the transfer analysis above, because the inversion is worth dwelling on. The same laboratory whose hosted endpoint you could not lawfully send a customer's correspondence to has handed you, at no cost, the means never to need it. Once the file is on your hardware there is no transfer, no processor, no jurisdiction question and nothing for anybody to compel, because there is no continuing relationship to compel anything through. The PRC National Intelligence Law has no more purchase on a set of weights running in your server room than the CLOUD Act has on a copy of PostgreSQL.


The objection to a Chinese model was never the model. It was the destination. Remove the destination and the objection goes with it.

Nor does this require a research budget. Summarising a thread, tagging by rule, translating a paragraph and generating embeddings are not frontier tasks. Small models handle them competently on a decent workstation, mid-sized ones on a single GPU in a server, and embedding models in particular are cheap enough to be an afterthought. For a great deal of routine mail work the quality gap against a hosted frontier model is narrow, and it is narrowing.


Four things worth knowing, two of which are usually got backwards:

  • Open weights are not open training data, and neither is anything else. You cannot audit what went into the model. That is true, and it is not a Chinese problem or an open-weight problem, it is universal: no laboratory in any jurisdiction publishes its corpus, and the American hosted providers disclose least of all. The difference is that open weights let you download the artefact, test it, benchmark it and run it in isolation on a machine with no route to the internet, which is a good deal more scrutiny than an API permits. On this axis open weights are the stronger position, not the weaker one.
  • Alignment reflects where it was trained, and not in the direction people expect. Every model carries the shaping of its origin, and on politically sensitive subjects you will see evasion. For business use, though, the difference that actually costs you time runs the other way entirely. Nothing refuses ordinary work with the frequency of the American hosted models: they decline to draft a firm letter before action, hesitate over a thread about a dispute, soften a deliberately blunt paragraph, and offer a short homily on tone while doing it. The open-weight models, and the Chinese ones notably, are far more tolerant and simply get on with the job you asked for. For a mail assistant, where the task is to draft what you told it to draft, that is not a minor consideration.
  • Read the actual licence. They vary far more than the headlines suggest, several are custom rather than OSI-approved, and at least one prominent release shipped under bespoke terms that a good many people assumed were MIT. Commercial use, redistribution and attribution conditions all differ.
  • A model file is a supply chain. Take weights from the publisher or a reputable mirror, verify the hashes, and pin the version you have tested, exactly as you would for any other dependency.

And one thing matters more than all four: running the model locally makes the processing private, it does not make the output right. A local model will invent a delivery date with precisely the same confidence as a hosted one. It fixes the data protection problem completely and the accountability problem not at all, which is why the human stays on the send button either way.


Where the Assistant Lives Matters

Assume you accept all of the above and want an assistant rather than an autopilot. The next question is whose software it runs inside, and that turns out to decide almost everything else.


With vendor-integrated AI in webmail or Outlook, the answers are fixed for you:

  • You do not choose the model. You get whichever model the vendor has commercial reasons to run this quarter, at whatever quality it is on the day, and it can change underneath you without notice or version pinning.
  • You do not choose the scope. The integration reads the mailbox, and often the calendar, the files and the chat history alongside it. Granularity is offered as a tenant-wide toggle, not a per-mailbox decision.
  • You do not choose the destination. The content of your mail leaves your estate to be processed, in a place and under a legal regime selected by somebody else. Contractual assurances about that processing are real, and they are also the vendor's to revise. You inherit the transfer risk assessment either way.
  • You do not choose the roadmap. Features appear, are enabled by default, are rebranded, are bundled into a higher licence tier, and occasionally are withdrawn. None of that is your decision, and all of it is your problem.

This is the ecosystem argument we have made before in 20260102, applied to the one system every business actually depends on. It is also the sovereignty argument from 20260115 and 20260303, and the hosting argument from 20240529: the more of your correspondence is processed somewhere you cannot point to on a map, the less of it is genuinely yours, and the harder it becomes to answer a data subject who asks a perfectly reasonable question about where their information went.


If the assistant is part of the mail service, you have not bought an assistant. You have extended the lease on the mail service.

Thunderbird and Betterbird: The Enterprise Client Nobody Talks About

Thunderbird is generally filed under "the free one", which does it a considerable disservice. Judged as an enterprise mail client rather than as a curiosity, it holds up rather well:

  • It is a client, not a tenancy. Standards-based IMAP, POP3 and SMTP, open calendar and contact protocols, and LDAP directory lookup. It talks to your mail server, whoever runs it, and it will still talk to the next one.
  • The mail is on a machine you own. Local profiles mean backup, archiving, retention and disclosure exercises are conducted against files under your control, on your terms.
  • Encryption is built in. OpenPGP and S/MIME ship as part of the client, not as a paid add-on or a premium security tier.
  • It can be centrally configured. Enterprise policies through a policies.json file allow account settings, restrictions and add-on control to be deployed across a fleet without touching each workstation.
  • The release cadence is sane. An extended support release with security updates through the year, rather than a rolling stream of interface changes nobody asked for.
  • It is open source, and that is a security control rather than a price. Thunderbird is published under the Mozilla Public Licence 2.0. The entire source is public, so anybody can read it, audit it, have a third party audit it, build it themselves and confirm the binary matches, or change it. You are not asked to believe a statement about what the client does with your mail; you can go and look. There is also no per-seat cost to model when headcount changes.
  • Exchange is reachable. Where a business is still on Exchange, add-ons such as Owl or TbSync bridge the gap, and native support has been steadily improving.

Betterbird is worth knowing about separately. It is a downstream build of Thunderbird, maintained by a former Thunderbird developer, tracking the same release trains but carrying fixes and features that have not made it upstream. For organisations that want the platform without waiting on an upstream release cycle for a defect that affects them today, it is a serious option, and it is the same profile format, so moving between the two is not a migration. It is also the open source argument demonstrated rather than asserted: because the source was public, somebody who disagreed with a decision could take it and fix it, and everybody else got the benefit. With a closed client, a defect that affects you is a support ticket and a hope.


Being honest about the trade: with a desktop client you own the deployment, the profile strategy, the backup and the update discipline. That work does not vanish, it moves to you. It is also precisely the work we do for clients every day, and it is a far smaller burden than the one you take on by handing your entire correspondence history to a platform whose commercial terms you do not set.


The Extension Ecosystem: Bolt On What You Want

The strongest argument for Thunderbird is not any single feature. It is that the feature set is not a fixed menu decided in another country by people who have never seen your business.


In a vendor-integrated client you get the tools the vendor built, in the order the vendor built them, and you campaign politely on a feedback forum for the rest. In Thunderbird you install what you actually need:

  • Owl or TbSync for Exchange, CalDAV and CardDAV connectivity.
  • Send Later for scheduled and recurring delivery.
  • Quicktext for templates and standard responses that your people write, rather than a model inventing them.
  • FiltaQuilla for filtering that goes well beyond the built-in rules.
  • QuickFolders and manual folder ordering for people who live in a fifty-folder mailbox.
  • ImportExportTools NG for bulk import, export and archiving, which is the tool you want to already have installed on the day somebody asks for six years of correspondence with one supplier.

None of that requires a licence upgrade, a tenant-wide change, or a conversation with an account manager. It requires a decision by you about what your people need.


One caution, because it is the same caution we gave in 20251128: an add-on is a supply chain. Review what it does, check that it is maintained, understand the permissions it asks for, pin the versions you have tested, and use the enterprise policy file to control what may be installed on managed machines. Freedom to bolt things on is only an advantage if somebody is deciding what gets bolted on. The saving grace is that these are overwhelmingly open source projects, so "review what it does" is a real instruction that somebody can actually carry out, rather than the polite fiction it becomes when the feature is a closed component of a platform you are obliged to accept whole.


GENA

Which brings us to our own contribution. GENA is an extension for Thunderbird and Betterbird that puts an AI assistant inside the mail client rather than beside it: summarise a thread, rewrite a draft, generate content from an inline instruction, digest an inbox, tag messages against your own rules, translate, strip the formatting that other clients insist on injecting, and search an entire mailbox by meaning rather than by keyword.


Four of its design decisions matter more than the feature list, and each of them is a direct answer to a problem set out above.


1. You Choose the Model

GENA does not hard-code a provider. You configure the endpoint, the headers, the request body and the path to the generated text in the response. That makes it work with Ollama or LM Studio on a workstation, vLLM on your own GPU, or a hosted service if that is what a particular job needs.


This is not a technical preference, it is five protections at once, and the first of them is the one that keeps a data protection officer in the room:

  • Compliance becomes a placement decision. Point GENA at a model on your own hardware and the entire apparatus described above simply does not engage: no third party, no disclosure by transmission, no Article 28 contract, no restricted transfer, no transfer risk assessment. Point it at a hosted service for a category of mail where you have done the work, and that is a documented decision rather than a default nobody made.
  • Confidentiality survives your NDAs. Legally privileged correspondence, HR matters, tender responses and commercially sensitive negotiation stay inside the building, so the confidentiality clause you signed is not quietly breached by a summarise button.
  • Capability is a per-job decision. Summarising an internal thread and drafting a formal response to a regulator do not need the same model, and there is no reason to pay for the second when doing the first.
  • Nobody changes the model underneath you. A model that behaves acceptably today can be deprecated, retuned or quietly re-routed tomorrow. When the endpoint is yours, that happens when you decide it happens.
  • There is an exit. Model choice that can be changed by editing a configuration field is not lock-in. It is procurement.

2. Semantic Search Over Your Own Mailbox

Keyword search fails constantly, and it fails for a boring reason: it matches the words in the message, and you are trying to remember the subject of the message. You are looking for the enquiry about bulk water supply. The email says "tanker deliveries" and never once uses the word water.


GENA embeds messages as vectors, indexes them in Qdrant, and searches by meaning. New mail is indexed as it arrives and deleted mail is pruned on the next scan. You ask a question in plain English, the relevant messages are retrieved, and the model answers with a supporting evidence table you can click straight through into the message itself.


The evidence table is the part that matters. An answer with no sources is a machine asking to be believed. An answer with the six messages it came from is a research assistant, and you remain the person who reads them and decides. And because the index is yours, the entire corporate memory of the mailbox does not have to be shipped into somebody else's search product to make it findable.


3. Semi-Manual Composition, Because Auto-Reply is Dangerous

GENA does not send email. It is a deliberate design constraint, not a missing feature.


Instead, you write the email, and where you want a paragraph generated you write an instruction on a line beginning with %. Press the GENA button and each instruction is replaced, in place, with generated content. The surrounding text is untouched, multiple instructions are processed in order, and nothing whatsoever leaves the client until you read the result and press send yourself.


In practice it looks like this:

Dear Mr Blobby,

Thank you for your email.

%explain why using Microsoft Outlook
for corporate email is a poor choice
and what we recommend instead

If you have any questions, please
let me know.

Regards
Rich

You have decided who to write to, what the email should say, what argument it should make, and what tone it should take. The model has done the typing for the one paragraph you did not want to type for the ninth time this month. That is a real saving, repeated many times a day, and it costs you nothing in accountability, because the accountability boundary is exactly where it always was: at the send button, with a person on the other side of it.


Compare that with an automatic responder. It also saves the typing. It also removes the person who would have noticed that the thread was hijacked, that the deadline stated in it was wrong, that the customer was not asking what the model thought they were asking, and that the paragraph promised a delivery date nobody in the business had agreed to.


  Assist (semi-manual) Take-over (automatic)
Who composes You, with the machine drafting on instruction The machine, from the thread
Who reads before sending A person, every time Nobody
Who is accountable The sender, knowingly The sender, unknowingly
Hallucinated commitment Caught before it leaves Sent, and binding
Hijacked or fraudulent thread Human scepticism applies Answered politely
Time actually saved Most of it Slightly more, at considerable cost

The efficiency difference between the two columns is small. The risk difference is the whole article.


4. You Can Read It

Everything above is a vendor making claims about its own software, which is precisely the category of statement this article has spent several thousand words advising you not to accept. So do not accept it. Check.


GENA ships as a .xpi, which is a zip file. Rename it, unpack it, and what is inside is the code that runs: plain JavaScript, HTML and CSS, not minified, not bundled, not compiled, with the developer's comments still in place. There is no build step between what we wrote and what your machine executes, and no separate source release you have to take on faith as matching the binary, because the shipped artefact is the source.


That turns every claim in this section into something testable in a text editor rather than something to be believed:

  • Want to know exactly what is sent to your chosen endpoint, and in what shape? Read the function that builds the request.
  • Want to know whether anything is sent anywhere else, to us or to anybody, on any schedule? Search for every network call in the extension. There are not many, and they are all in front of you.
  • Want to know what the semantic index stores and where? Read it. The Qdrant instance is one you point it at.
  • Want to satisfy yourself that it genuinely cannot send an email on its own? That is a question about which APIs the extension requests and calls, and both are stated in the manifest and visible in the code.

Your own security team can review it before it goes anywhere near a mailbox. Your auditor can be handed the file. If it does something you dislike, you can change it, and having reviewed a version you can pin that version across the fleet through the same enterprise policy file that controls everything else. None of that requires our cooperation, our permission, or our continued existence as a company.


You should not have to trust something with the whole of your correspondence. You should be able to read it.

Now ask the same questions of an assistant built into a mail platform. What is transmitted, to which region, under what retention, and does that change next quarter? The best available answer is a documentation page describing intended behaviour, written by the party with the least interest in your reading it closely. That is a description of a promise. The code is a description of a mechanism, and only one of the two can be checked.


What an Inbox AI Policy Should Actually Say

If you are introducing AI into mail, this is the short version of the policy, and it fits on one side of paper:

  • No message leaves the organisation without a human having read it. No exceptions for out-of-hours, holidays or high volume. If the volume is the problem, the answer is staffing or templates, not an unattended model.
  • Summaries are a reading aid, not a record. Nobody acts on a summary of a contractual thread without opening the thread.
  • Define where mail may be processed, by category. Name the classes of correspondence that may only ever reach an internal model, and enforce it by configuration rather than by trusting people to remember on a busy afternoon.
  • Do the paperwork before the pilot, not after it. A legitimate interests assessment, a DPIA where the processing warrants one, the provider added to your record of processing and your sub-processor list, an Article 28 contract in place, and your privacy notice updated to say what actually happens. If that cannot be completed, the feature is not ready to be switched on.
  • No shadow AI, and name the permitted tiers. Staff pasting a customer thread into a consumer chatbot is the same transfer, into the one tier that trains on it, with no contract at all. Specify the exact services and tiers that may be used, block the rest, and give people a sanctioned tool good enough that they do not go looking.
  • Automatic classification may sort, never delete. A tagging mistake is an inconvenience; a deletion or an auto-archive is a missed notice period.
  • Anything touching money follows the verification process regardless. No summary, tag or draft alters the requirement to verify bank details against an independent source.
  • Prefer software you can read. Where a tool is going to process the whole of your correspondence, open source is not an ideological preference, it is the only way anybody outside the vendor can establish what it does. Make auditability a procurement criterion and require that the version deployed is the version somebody reviewed.
  • Log what the assistant was asked and what it produced. When somebody later asks why an email said what it said, you want a record rather than a theory.
  • Review the drafting register. A model trained largely on American corporate English will quietly Americanise your correspondence, and clients notice.

Summary

AI in the inbox is genuinely useful, and the useful part is not the part being marketed hardest. Summarising, searching, drafting and tagging save real time for real people. Sending on your behalf saves a few more seconds and hands a statistical text generator the authority to speak for your company in writing.


And remember whose data you are being invited to be generous with. A mailbox is not your information, it is thousands of other people's information, held by you because they had to send it somewhere. They did not consent to it being transmitted abroad for a model to process, they were never asked, and in most cases they could not sensibly have been asked. That makes the choice of where processing happens a governance decision at board level, not a preference buried in a settings panel.


Weigh it accordingly, because this is one of the few decisions in IT that cannot be reversed. A bad server can be replaced, a bad supplier can be exited, a bad contract can be renegotiated. Text that has been absorbed into somebody else's model weights stays there, and no amount of later diligence gets it back.


Keep the assistant on the assist side of the line. Run it in a client you control rather than in a service that controls you. Choose your own model, and favour a local one, so that confidentiality, jurisdiction, capability and cost all remain your decisions. Index your own mail so you can find it by meaning without surrendering it to somebody else's search index. Prefer open source at every layer, the client, the extension and where you can the model, so that "what does this actually do with my mail" is a question answered by reading rather than by asking. And keep a human on the send button, because that is not a limitation of the technology, it is the entire point.


GENA for Thunderbird and Betterbird  GEN Email Services  Data Protection

About GEN

GEN has been working with machine learning since 1999, long before "AI" was a marketing word. We run our own compute clusters, train and fine-tune our own models, and write our own post-processing and guardrail layers rather than wrapping someone else's API and hoping for the best. We provide commercial AI solutions to enterprise customers globally, with confidentiality and security as a core part of what makes our systems valuable. Client prompts and embeddings are never stored, indexed, shared, captured, or used for training. We believe large language models are most useful when they are applied carefully and appropriately to real-world problems, not treated as something to stuff into a workflow and hope for the best.


             8 Votes  
100% Human Generated

×

--- This content is not legal or financial advice & Solely the opinions of the author ---

Contact Us