Comparing AI tools? Think about where your data will end up
If you’re giving AI access to sensitive information, it’s worth considering who you’re sharing the data with.
In our current AI-fueled world, it seems every organisation is trying to find more ways to offload data automation to LLMs. Offloading your sensitive data workloads to foreign companies with murky privacy policies isn’t something companies should be doing lightly.
Who’s using your data?
Widespread AI usage makes data privacy and ownership concerns more tangible and pertinent than ever. Especially considering frontier model providers have a habit of ensuring that whatever you share with them, they get to keep.
If you’re putting lots of data into an LLM, you’ll need to know what’s being done with it on the other end. Frontier models offered by companies like Claude and OpenAI are a bit of a black box—it’s impossible to know what’s really being done with your data. Sure, these providers let you tweak privacy settings, but they’re still retaining your inputs and outputs for up to 2 years. Even conversations deleted from the Claude back-end still sit in their storage system for up to another 30 days. That’s not particularly private.
It becomes a near impossibility to know where the data will end up. Overseas companies don’t have the same data usage obligations that New Zealand companies have.
On the free and lower-priced tiers, their policies outline exactly why you shouldn’t trust them. On the pricier plans, you’re trusting international vendors with a record of leaking user inputs and publishing private AI conversations to handle your data responsibly.
Data fed to these offshore LLMs often becomes property of the AI providers to use as they please. Unless your provider is saying otherwise, anything thrown at AI models can be logged and used to retrain future models. That puts you in a precarious position, where you’ll have no idea when and where data you’ve fed it will resurface elsewhere.
Competitive jeopardy
Anything that gets used for training can be reused as answers or other outputs later. It’s not just the devs at OpenAI and Anthropic you can see your inputs. That data could potentially reappear to anyone who uses the LLM. Documents summarised by generative AI aren’t always deleted, and can even turn up in another user’s output.
If your organisation is giving AI access to sensitive info, you could be unwillingly sharing information your competitors don’t know about. Letting LLMs regurgitate confidential data all across the world isn’t something you should simply accept.
Organisations working with other people’s information can’t be flippantly feeding that data into LLMs without considering the privacy implications—you don’t really know what data will remain private, and you could be violating New Zealand privacy laws.
At best it’s a competitive risk. At worst it’s literally illegal behaviour.
Local risks, and NZ Privacy law
Breaching an individual or an entity’s privacy rights is serious business.
New Zealand’s Privacy Act 2020 is the country’s most conclusive set of rules concerning information privacy. It’s not designed specifically for AI use, but it governs how organisations and businesses can collect, store, use and share personal information. It’s a lengthy document full of legalese, but there are a few pointers in it that are particularly relevant for businesses using AI.
Rather than untangling our way through it, we may as well cite the Privacy Commissioner’s own expectations around AI usage in the country and map potential risks against the Privacy Act’s Information Privacy Principles.
There’s plenty of stuff that you can do with AI, but there are good reasons not to do it with Claude or ChatGPT.
One of those principles states that organisations must put safeguards in place to prevent loss, misuse or disclosure of personal information. Pouring data into LLMs run by companies who train models on your inputs and outputs seems to defy that principle.
And then there are other principles all about collecting information from people, and disclosing to them how the data will be used. If you’re directly collecting data, you’re supposed to take reasonable steps to make sure individuals know their data is being collected, and inform them who will be receiving it. You’re also supposed to let individuals know how the data will be used, and if they have the option to opt-out of the data collection.
When you’re using frontier LLMs, you’re losing control of how the data is being used and whose hands it ends up in. It becomes a near impossibility to know where the data will end up. Overseas companies don’t have the same data usage obligations that New Zealand companies have. If your organisation is the one feeding AI companies confidential data, how would you justify it to the Privacy Commissioner?
Sectors where it (really) matters
For every organisation, data privacy is something to weigh up and think about. For organisations in heavily regulated industries, it’s a legal obligation to consider how AI handles their data.
Three notable examples are health, legal, and financial services; each operating with sector specific privacy rules and guidelines on top of the Information Privacy Principles.
Health
Hospitals, pharmacies and healthcare providers have to take very intentional care of confidential data to adhere to the Health Information Privacy Code 2020. Uploading patient records to an international AI provider could easily fly directly in the face of that. It’s easy to imagine the negative consequences if patient data that’s been fed into a public AI model and stored overseas goes on to resurface again elsewhere.
Law
Legal professionals bound by client-attorney privilege and strict confidentiality agreements shouldn’t risk feeding case files, contacts or client disclosures into public LLMs.
Don’t take it from us, take it from the NZ Law Society’s Generative AI Guidance.
“Inputting client details and legally privileged material into a publicly accessible/external Gen AI tool may also give rise to a breach of privilege and confidentiality obligations. At a minimum, lawyers need to consider whether client consent should be sought for use of their data.”
It’s not just the individual’s rights that are worth considering, the international location of the servers hosting the LLMs also matters.
“Be aware that the AI provider may well be able to see your input data and the outputs. This can create a privacy risk (in addition to concerns about confidentiality and privilege). The data inputted may also be transferred out of New Zealand to AI companies located overseas. This has implications under the Privacy Act and lawyers should have regard to Information Privacy Principle 12.”
Finance
Banks, insurers and other financial institutions operate under stringent frameworks from regulatory bodies like the Financial Markets Authority, who are currently undertaking a review on AI usage in the sector as we write this.
What we can say for sure, is that sharing sensitive customer records, transaction histories or assessments with foreign-hosted AI vendors potentially exposes institutions to risk of compliance breaches and potential data leaks.
If AI workloads are going to be happening—as with health and law—one way to guarantee the data never goes offshore or sees the light of day again is to only use local providers with zero data retention.
Government holds itself to this bar
At least in New Zealand, public sector agencies are bound by cloud and data directives holding them to high standards of data privacy and sovereignty (we’ll get more into this later). It’s understandable, given the huge public trust the population gives them to handle data representing New Zealand’s ~5 million civilians.
Cabinet mandated cloud risk assessments and Māori data governance principles mean Government organisations assess cloud services on a case-by-case basis. Unsurprisingly, the risks associated with international cloud services providers often demand locally-hosted and owned infrastructure that keeps foreign legislation like the US CLOUD act out of the conversation. That excludes big American AI providers like OpenAI, Google and Anthropic—even if they don’t intend to use the data for training, they still hold on to it.
The same level of scrutiny is expected from AI providers.
Be careful
Since feeding sensitive data to frontier models introduces a lot of legal and competitive risks, you’re often left with no choice but to limit what you share with AI services. This can be a bottleneck, preventing organisations from fully capitalising on the potential of AI automation.
Of course, there will always be work that’s suitable for frontier models. Claude models are still going to be a strong choice for deep reasoning work, and they remain at least a few months ahead of any competition. If privacy isn’t a concern, and you’re primarily in the market for raw capability and reasoning, then frontier models are still a good way to go.
When you’re working with protected data, whether it’s trade secrets or other sensitive information, you’ll need to think about the contents of what you’re offloading to the frontier models.
Or you’ll need to find a more secure way to offload that workload.
Frontier model providers have a habit of ensuring that whatever you share with them, they get to keep.
All of these privacy concerns add up to one outcome: there’s plenty of stuff that you can do with AI, but there are good reasons not to do it with Claude or ChatGPT. By taking the right amount of care around privacy, you can end up limiting your access to the tech that’s most likely to accelerate your business today. If only there was a way to access AI that didn’t incur privacy downsides.... Oh, hey look!
AI without data retention
The cleanest way to navigate sensitive data concerns is to only use AI providers with transparent zero data retention policies. Frontier models from Anthropic and OpenAI don’t give you that. You’ll want to find providers running customisable open-weight models, which can guarantee that no data going in goes anywhere else.
And that’s exactly what we’re doing here at SiteHost. Our recently launched AI Platform is a New Zealand-hosted, locally-owned, zero-data retention service for running AI workloads.
By design, these open-weight models have completely customisable parameters. The way we’ve got them set up, there’s zero data retention at any level. No model training, no data retention.
AI Platform mitigates data privacy concerns by not collecting any of your data in the first place. You can feed data into the platform without ever worrying that it will go elsewhere.