For most of the history of school software, a grade was computed in your own database. A mark was a sum, a lookup, a rule. The student's work sat on infrastructure you chose, in a region you could name, and the privacy conversation with a buyer was short.

Once an LLM marks the answer, a student's free text leaves that building. It is sent over an API to a model that may run in another country, operated by a company that is now a party to the processing. Three things have changed at once. Personal data crosses a border. A new organisation handles it. And somewhere outside your control, a copy may be kept.

This is an architecture article, not legal advice. The regimes involved differ and change, and a procurement committee will have counsel. What I can do is set out the technical decisions that decide whether you can give a straight answer when that committee asks.

Start by mapping what actually leaves

Teams usually describe the flow as "we send the answer and the rubric to the model". That is one of at least four copies of the data, and the other three are where the trouble is.

The first is the request payload: the student's response, any identifiers you included, the rubric, and any context such as earlier answers. The second is whatever the model provider keeps. Many providers retain inputs for a period for abuse monitoring unless you hold an agreement that disables it, and that retention is separate from any statement about training. The third is your own telemetry. Traces, logs, error captures and queue payloads routinely hold the full prompt, and they live wherever your observability vendor does. The fourth is your evaluation and debugging data: the golden sets, the exports into notebooks, the sample a developer pasted into a ticket.

Residency work that stops at the first copy is cosmetic. Draw all four on one page before you do anything else, and mark where each one physically sits and who can read it.

Diagram of at least four copies made from one graded answer. The request payload, the one teams describe, holds the response, identifiers, rubric and context. The other three are marked as where the trouble is: what the provider keeps, often for abuse monitoring; your telemetry, such as traces, logs and queue payloads; and evaluation and debugging data. One graded answer, at least four copies 1 Request payload the one teams describe the response, identifiers, rubric and context The other three are where the trouble is 2 What the provider keeps often retained for abuse monitoring 3 Your telemetry traces, logs, error captures, queue payloads 4 Evaluation and debugging data golden sets, notebooks, samples in tickets Mark where each sits and who can read it.
Residency work that stops at the first copy is cosmetic.

The basis for processing, and why minors change the conversation

In most institutional deployments the school or university is the data controller and you are a processor acting on its instructions. That matters, because the legal basis for processing, and any consent from a guardian, sits with the institution, and your obligations are largely set by the processing agreement and by what you can show you did.

Minors raise the stakes. India's Digital Personal Data Protection Act, 2023 requires verifiable consent from a parent or guardian before processing a child's data, and restricts tracking and behavioural monitoring aimed at children, with provisions for educational institutions that should be read in the rules rather than assumed. The GDPR gives children specific protection and treats cross-border transfers under Chapter V, which in practice means standard contractual clauses and a documented transfer assessment. In the United States, FERPA treats a vendor as a school official only under direct institutional control, and COPPA applies to services directed at under-13s. The UAE and Saudi Arabia each have a federal personal data protection law with its own transfer conditions.

You do not need to master all of them. You need to know, per contract, which one the buyer is bound by, and to have an architecture flexible enough to satisfy the strictest realistic one without a rewrite.

Two architectures, and what each one actually buys you

There are two ways to reduce what crosses the border, and they are routinely confused.

A regional inference endpoint keeps the data in a chosen geography. When you assess one, ask three separate questions, because a "regional" label can answer the first and not the others. Where does inference run? Where does data sit at rest, including any cached or queued copies? And what does the provider retain, for how long, and can that be disabled by contract?

Then ask the question teams forget, which is failover. A service configured for one region that quietly falls back to another under load or outage has broken its residency promise at exactly the moment nobody is watching. Residency-sensitive traffic should fail closed: if the in-region endpoint is down, the request errors or queues, and an operator decides. The inconvenience is real, and it is the cost of the claim.

A line of pale canisters waiting on one plinth in front of a bridge to a second, empty plinth, with the bridge's entrance shut by an orange gate, showing residency-sensitive traffic that fails closed instead of falling back to another region.
Failover is where a residency promise quietly breaks. Residency-sensitive traffic should fail closed: the request errors or queues, and an operator decides.

Redaction or pseudonymisation at the boundary strips or tokenises names, roll numbers and other identifiers before the call, and restores them on return. It is cheap, it keeps working with any provider, and it reduces the harm if a copy leaks. It has three limits. Free text identifies people in ways a field-level scrubber misses: a student writing about their own school, a named teacher, a locality, a handwritten answer with a name in the margin that OCR faithfully transcribes. Named-entity recall is high on clean English and noticeably worse on transliterated names and mixed-language text. And it does not change the legal status. Under the GDPR, data that can reasonably be re-identified remains personal data (Recital 26), so redaction lowers risk without taking the data out of scope.

That last point decides which one a buyer will accept. A procurement requirement framed as "personal data shall be processed within the country" is not met by redaction. It is met by in-country inference, which means a regional cloud presence or a self-hosted open-weight model. The cost comparison there is not as simple as it looks. Per-token pricing is cheap at low volume and a dedicated GPU is cheap at sustained volume, and an exam calendar is neither: it is a handful of intense peaks on a quiet baseline. Utilisation across the whole year, not the peak, decides whether self-hosting pays.

My working rule is to build both layers regardless. Redact by default so that the model never receives identifiers it does not need, and make the endpoint region a routing decision per tenant rather than a deployment constant. That way a contract that demands in-country processing is a configuration change, and one that does not still gets the benefit of the redaction.

Diagram of the two layers on every request. Names, roll numbers and other identifiers are tokenised at the boundary, which lowers risk but leaves legal status unchanged. The request is then routed by tenant: a contract that demands in-country processing gets in-country inference, and one that does not still gets the redaction. Identifiers are restored on return. Redact by default, route by tenant Redact at the boundary names, roll numbers tokenised Lowers risk; legal status unchanged Route by tenant not a deployment constant Makes in-country a configuration change In-country contract in-country inference No in-country clause still gets the redaction Identifiers restored on return
Redaction reduces risk without changing legal status, so it will not satisfy an in-country requirement. Making the region a per-tenant routing decision turns that requirement into a configuration change.

Measure what redaction costs you. Run your regression set both ways, with identifiers present and removed, and compare agreement with the human marks. It is usually small, but it is not zero on essays that refer to people, and you want that number before a customer asks.

Retention: the record you must keep and the data you must not

Assessment has a property most software lacks. A disputed result has to be reconstructable months later, which pushes you towards retaining the answer, the rubric version, the model version, the score and the rationale. Privacy law pushes the other way, towards minimisation and erasure on request.

The way out is separation, not compromise. Keep the evidential grading record in the buyer's region under a defined retention schedule. Keep the mapping from that record to a named person behind a separate key, held apart. An erasure request then removes or destroys the mapping, and the evidential record stops being attributable to an individual without losing its integrity as a record. Whether a given regime accepts that for a given dataset is a legal question, but it is the design that gives counsel something to say yes to. The storage design that makes a non-repeatable generation reconstructable deserves its own treatment, and I will give it one.

Diagram of two stores held apart. The evidential record, with the answer, score, rationale and the rubric and model versions, stays in the buyer's region on a retention schedule. The mapping from that record to a named person sits behind a separate key. An erasure request removes the mapping, so the record stays whole but is no longer attributable. Erasure removes the mapping, not the record Evidential record answer, score, rationale rubric and model versions in the buyer's region, on a retention schedule Identity mapping links the record to a named person separate key, held apart Erasure request removes or destroys it Record stays whole no longer attributable Whether a regime accepts it is a legal question.
Keep the evidential record in the buyer's region and the identity mapping apart, so that erasure and reconstruction can both be true.

On the provider side, do not rely on their retention at all. The record you keep is your own, in your own region, and the provider's job is to forget.

Sub-processors: your supply chain is now part of the answer

A buyer is entitled to know who touches the data. The list usually runs longer than teams expect: the model provider, the cloud, the OCR or speech service, the observability vendor, the email or SMS gateway, the error tracker. Every one of those is a sub-processor in the plain sense, and each needs a place on a disclosed list with a stated purpose and location.

Two operational points. Notice of change is normally a contract term, not a statutory minimum, and varies by agreement, so decide what you can honestly commit to. And a model upgrade that moves you to a different provider is a sub-processor change, so version pinning and a controlled migration are compliance controls as well as quality controls.

What the questionnaire actually asks

The questions that recur are not exotic. Where is the data stored and where is it processed? Who has access, and is that access logged? Is customer data used to train any model? How long is it retained, and how is it deleted at contract end? What is the breach notification window? Who are the sub-processors? Can the customer audit, or rely on an independent report? How are encryption keys managed?

Each of these is a claim, and each should map to something you can show: a configuration, a diagram, a log, a contract clause. The failure I see most often is not a missing control. It is an answer written by someone selling, in good faith, that engineering cannot back. Keep one standing answers document, owned by engineering, reviewed by counsel, and updated whenever the architecture changes. Sales copies from it and does not paraphrase it.

Diagram of a chain from what you can show, such as a configuration, a diagram, a log or a contract clause, to one standing answers document, owned by engineering, reviewed by counsel and updated whenever the architecture changes, and from there to questionnaire answers that sales copies without paraphrasing. One source of answers, owned by engineering What you can show configuration, diagram, log, contract clause updated whenever the architecture changes Standing answers document owned by engineering, reviewed by counsel sales copies it and does not paraphrase it Questionnaire answers each one a claim engineering can back
Every questionnaire answer is a claim. Keep one standing answers document that engineering owns and counsel reviews, and let sales quote it but not rewrite it.

What I got wrong

I have fixed an inference path so that student text stayed in the intended region and then found the full prompts still sitting in the observability stack in another one. The model call was compliant. The trace of the model call was not.

And I have seen a questionnaire answer say that data stays in-region while a failover route, added later for availability, sent overflow to a different region. Each team was right about its own part, and the answer was wrong.

The short version

Draw the four copies of the data, not one. Treat the regional endpoint as three questions and fail closed on failover. Redact by default, and understand that redaction reduces risk without changing legal status, so it will not satisfy an in-country requirement. Keep the evidential record in the buyer's region and the identity mapping apart, so erasure and reconstruction can both be true. List every sub-processor, pin your model versions, and keep one engineering-owned set of answers that a sales team can quote but not rewrite.

None of this slows a good team down. It is much cheaper to build in the first quarter than to explain in the second year.

Illustrations generated with AI.