Where does your data go when you use cloud AI?
The content you send is processed by the provider. Your prompt, attached documents, and associated metadata can leave your own systems. Connected apps, logs, backups, and support access may create additional routes. Draw those routes before deciding whether a product fits the workflow.
Location is only one part of the answer. You also need to know who can use the data, how long it is kept, and how it is removed. A promise about model training does not answer those questions.
Does the provider train on your data?
Check the exact product, tier, settings, and agreement. For example, OpenAI says business and API data is not used for training by default, with exceptions for data you explicitly opt to share. That does not mean the service processes no data or retains nothing.
Record the training terms, retention settings, processing location, subprocessors, and connected-app permissions together. Give staff an approved tool and a short rule about which information may go into it. Do not assume a personal account has the terms your business needs.
What changes when the data is regulated?
Your legal, contractual, and operational obligations become part of the design. Encryption and a no-training clause are controls, not a complete compliance decision.
For healthcare, HHS permits cloud processing of ePHI with the required BAA and compliance safeguards. HIPAA does not simply require a server in the practice. Have the responsible reviewer assess the data flow, agreements, access, retention, and incident responsibilities before patient data is connected.
What does on-premise AI actually mean?
On-premise AI runs on hardware at your premises. A model in a private cloud tenancy is hosted elsewhere, even if the infrastructure is dedicated to you. Either design still needs a map of external connections, backups, telemetry, updates, and support access.
MARCUS, built for B:Side Capital where our founder is CEO, processes borrower documents locally. Selected tasks can use external reasoning on filtered text. Consequential actions require human approval. That is a specific architecture, not a claim that all processing stays inside the building. The security page describes its controls and their limits.
What is a hybrid AI deployment?
A hybrid deployment keeps some processing local and sends permitted material to external services for selected tasks. A privacy filter can detect supported identifiers and replace them before that transfer.
Filtering does not prove that the remaining text is anonymous. Presidio documents that automated detection can miss sensitive information. Context can also identify someone. Decide what may leave, test representative and difficult cases, and define what happens when the filter or the external service fails.
Which deployment fits which data?
Use this as a starting checklist. The most sensitive input, contractual restriction, and required control can change the answer.
| Data | Possible starting point | Check before use |
|---|---|---|
| Public material | An approved cloud tool | Usage rights, accuracy, and whether attachments contain nonpublic data |
| Internal business material | A business or API service with suitable terms | Access, training terms, retention, location, and connected apps |
| Customer identifiers | A reviewed cloud, hybrid, or local design | Permitted processing, minimization, filter limits, and contractual duties |
| Regulated or restricted records | A design approved against the actual requirements | Required agreements, risk review, access, logs, retention, and incident handling |
A local server is not a compliance certificate. A cloud service is not automatically unsuitable. The documented data flow and responsibilities decide.
Does a 20-person business need on-premise AI?
Headcount is not the test. Start with the data that cannot be sent to a third party, the terms you have agreed to, and the workload the model must handle. If an approved cloud service meets those needs, compare it with the cost and responsibility of operating hardware.
For a local option, budget hardware, administration, backups, security updates, electricity, and model testing. For cloud or hybrid, budget usage, tools, access controls, and oversight. Our running-cost guide separates those expenses.
What should you ask before signing?
- What exact records and fields does this workflow use?
- Where does each copy go, including prompts, logs, and backups?
- Who can access it, and which providers or subprocessors handle it?
- What are the training, retention, deletion, and exit terms?
- Which actions need approval, what failures stop the workflow, and who investigates?
Keep the answers with the scope. If the data path is unclear, that is useful discovery work for an AI Readiness Audit. For an initial discussion, bring the type of data and the workflow to the free 30-minute assessment; use examples without sensitive records until the handling arrangements are agreed.