Booking Q4 deliverystart with a free workflow plan Denver · Phoenix · Remote
The Field Guide / 09

Where does your AI data actually go?

Most owners sign the vendor agreement without reading the data clause. This guide is that clause, translated — cloud, on-premise, and the middle path, by data type.

Field GuideNo. 09
Reading time4 min
Paths comparedCloud · hybrid · on-prem
On-prem proofMARCUS · 14 agents

Where does your data go when you use cloud AI?

The content you send is processed by the provider. Your prompt, attached documents, and associated metadata can leave your own systems. Connected apps, logs, backups, and support access may create additional routes. Draw those routes before deciding whether a product fits the workflow.

Location is only one part of the answer. You also need to know who can use the data, how long it is kept, and how it is removed. A promise about model training does not answer those questions.

Does the provider train on your data?

Check the exact product, tier, settings, and agreement. For example, OpenAI says business and API data is not used for training by default, with exceptions for data you explicitly opt to share. That does not mean the service processes no data or retains nothing.

Record the training terms, retention settings, processing location, subprocessors, and connected-app permissions together. Give staff an approved tool and a short rule about which information may go into it. Do not assume a personal account has the terms your business needs.

What changes when the data is regulated?

Your legal, contractual, and operational obligations become part of the design. Encryption and a no-training clause are controls, not a complete compliance decision.

For healthcare, HHS permits cloud processing of ePHI with the required BAA and compliance safeguards. HIPAA does not simply require a server in the practice. Have the responsible reviewer assess the data flow, agreements, access, retention, and incident responsibilities before patient data is connected.

What does on-premise AI actually mean?

On-premise AI runs on hardware at your premises. A model in a private cloud tenancy is hosted elsewhere, even if the infrastructure is dedicated to you. Either design still needs a map of external connections, backups, telemetry, updates, and support access.

MARCUS, built for B:Side Capital where our founder is CEO, processes borrower documents locally. Selected tasks can use external reasoning on filtered text. Consequential actions require human approval. That is a specific architecture, not a claim that all processing stays inside the building. The security page describes its controls and their limits.

What is a hybrid AI deployment?

A hybrid deployment keeps some processing local and sends permitted material to external services for selected tasks. A privacy filter can detect supported identifiers and replace them before that transfer.

Filtering does not prove that the remaining text is anonymous. Presidio documents that automated detection can miss sensitive information. Context can also identify someone. Decide what may leave, test representative and difficult cases, and define what happens when the filter or the external service fails.

Which deployment fits which data?

Use this as a starting checklist. The most sensitive input, contractual restriction, and required control can change the answer.

DataPossible starting pointCheck before use
Public materialAn approved cloud toolUsage rights, accuracy, and whether attachments contain nonpublic data
Internal business materialA business or API service with suitable termsAccess, training terms, retention, location, and connected apps
Customer identifiersA reviewed cloud, hybrid, or local designPermitted processing, minimization, filter limits, and contractual duties
Regulated or restricted recordsA design approved against the actual requirementsRequired agreements, risk review, access, logs, retention, and incident handling

A local server is not a compliance certificate. A cloud service is not automatically unsuitable. The documented data flow and responsibilities decide.

Does a 20-person business need on-premise AI?

Headcount is not the test. Start with the data that cannot be sent to a third party, the terms you have agreed to, and the workload the model must handle. If an approved cloud service meets those needs, compare it with the cost and responsibility of operating hardware.

For a local option, budget hardware, administration, backups, security updates, electricity, and model testing. For cloud or hybrid, budget usage, tools, access controls, and oversight. Our running-cost guide separates those expenses.

What should you ask before signing?

  1. What exact records and fields does this workflow use?
  2. Where does each copy go, including prompts, logs, and backups?
  3. Who can access it, and which providers or subprocessors handle it?
  4. What are the training, retention, deletion, and exit terms?
  5. Which actions need approval, what failures stop the workflow, and who investigates?

Keep the answers with the scope. If the data path is unclear, that is useful discovery work for an AI Readiness Audit. For an initial discussion, bring the type of data and the workflow to the free 30-minute assessment; use examples without sensitive records until the handling arrangements are agreed.

Fair questions

Where AI data goes.

Want the numbers?

The full price list is published — audits, sprints, managed services, all on one page.

Read the price list →
01Does ChatGPT use my business data to train its models?+

OpenAI states that its business products and API do not use business data for training by default. Consumer settings and other providers have different terms. Check the exact product, retention settings, connected tools, and current agreement before using sensitive records.

02Does a small business need on-premise AI?+

The answer depends on the records, contracts, access requirements, and operating budget. Local processing can help restrict data routes, while approved cloud services may fit other workflows. Neither deployment choice alone makes a system compliant or private.

03How do you use AI on regulated data like health or lending records?+

MARCUS processes borrower documents locally on B:Side-owned hardware. Selected tasks can use external models on filtered text. The agreed data boundary determines which processing routes are permitted. Presidio, originally developed at Microsoft, detects supported personal identifiers for removal before model processing. Automated detection can miss information; filtering is one control alongside routing restrictions, access decisions, and tests using known identifiers.

04What is the middle path between cloud AI and on-premise AI?+

A hybrid deployment can process documents locally and send selected, filtered text to an external model. Identifier detection can miss sensitive information; decide what may leave, test the filter, and prohibit external processing where the risk or agreement requires it.

Start here

Not sure which deployment your data needs?

Tell us what one workflow touches: patient records, borrower files, or ordinary business documents. We will discuss the requirements and the next step in a free 30-minute assessment. We reply within 24 hours.

We reply within 24 hours. A fixed quote follows an agreed scope, before paid work begins.

— Christopher Myers, Founder