At Oracle AI World 2025, Larry Ellison said something that should make every organisation pause.

“We took all of our customer data, and we vectorized it.”

For most people, that sentence needs translation.

To vectorise data means converting information into a mathematical format that an AI system can search and compare by meaning, not just by keywords. It is a bit like turning documents, records and customer history into coordinates on a map. Once that map exists, the AI can quickly find related information, spot patterns, connect similar cases and retrieve the most relevant material for a question or task.

So when private customer data is vectorised, it does not simply sit in a database waiting for a human to open a record. It becomes searchable and usable by AI.

That is the real shift.

For the past few years, much of the public debate about artificial intelligence has focused on training data. What was scraped from the internet? Were creators asked for permission? Did personal information end up inside large language models? Can people opt out?

Those questions still matter. But they are no longer enough.

The next phase of AI is not just about training models on more public information. It is about connecting AI systems to private data: customer records, internal documents, emails, case notes, health records, financial information, regulatory files, employee data, research, operational systems and cloud storage.

In this new phase, the AI does not necessarily need to be trained on private data to use it.

That distinction is important.

A vendor may say, “We do not train the model on your data.” That may be true. But it does not mean the AI cannot retrieve the data, reason over it, summarise it, infer from it, combine it with other information or recommend actions based on it.

That is where the risk changes.

Ellison described Oracle’s AI database and AI data platform as a way to make private data available to AI models for reasoning. The technical approach is often called retrieval-augmented generation, or RAG. In simple terms, RAG allows an AI system to search private information, retrieve the most relevant pieces, and provide them to the model as context for a specific answer or task.

The data may remain inside the organisation’s systems. The model may not permanently learn from it. The private information may not be used to train a public model.

But the AI can still use it.

And once AI can use private data, the privacy and governance questions become much bigger.

Ellison’s example was not hypothetical. He said Oracle took its customer data, vectorised it and used RAG to make it available to AI models. He then described asking which Oracle customers were likely to buy another Oracle product in the next six months, what product they might buy, and how Oracle could send those customers targeted messages using relevant customer references.

From a commercial perspective, this is powerful.

From a governance perspective, it should make us stop and think.

The uncomfortable question is not whether Oracle technically had the data.

It did.

The uncomfortable question is whether customers understood that their data could be converted into an AI-searchable format and used to predict their future behaviour.

Were they meaningfully informed? Could they opt out? Or did private customer data simply become AI fuel because it was already sitting in the system?

That question matters because this is not just a technical change. It is a change in use.

Data that may have been collected for account management, service delivery, sales, contracts or support can become part of an AI reasoning layer. Once that happens, the data is no longer just being stored or queried. It is being used to predict, infer, personalise and potentially trigger action.

And if that can happen with customer data, what does it mean for more sensitive datasets?

What does it mean for health records, employee files, complaints, grants, investigations, regulatory information, citizen service records or data collected under legislation?

This is especially important for government.

Public sector organisations hold some of the most sensitive information in society. Citizens and businesses often provide that information because they are required to, because they need a service, because they are applying for support, or because they are subject to regulation.

The power relationship is different.

The duty of care is higher.

A citizen may provide information for a licence, grant, inspection, medical service, biosecurity process or compliance matter. That does not automatically mean they expect the information to be used by an AI system to generate predictions, risk profiles, recommendations or automated actions.

A business may provide information to government for reporting, regulation, funding or service delivery. That does not automatically mean it expects the data to become part of an AI reasoning layer that can be queried, combined and interpreted in new ways.

The issue is not whether AI can create value from this data. It clearly can.

The issue is whether usefulness is being mistaken for permission.

This is where the phrase “we do not train on your data” can become dangerously incomplete. It answers one question, but not the most important one.

The more important question is:

Can my data be used by the AI to create responses, predictions or recommendations for someone else?

That is the real concern.

In a retrieval-based AI system, the model has to access data to produce a useful result. If I ask an AI assistant to summarise my own records, then of course it needs access to those records. That is the point.

The risk emerges when private data is used beyond the person, organisation or purpose it was provided for.

Could my customer history be used to predict what another team should sell me?

Could my health information be used to recommend a funding, insurance or service decision?

Could my business records be used to generate a risk profile?

Could my complaint, grant application, inspection result or regulatory history be used to produce insights for someone else without my knowledge?

Could sensitive information about one person become part of an answer given to another person?

These are the questions that matter.

The issue is not simply whether the AI can access data. In many cases, it must access data to be useful. The issue is whether that access is limited to the right person, the right purpose and the right context.

For highly sensitive data, organisations need clear boundaries.

My data should not be used to answer another person’s question unless there is a lawful, authorised and transparent reason.

My data should not be used to create predictions or profiles about me unless that use is fair, explainable and expected.

My data should not be used to trigger decisions or actions without appropriate human oversight.

My data should not quietly become part of a broader AI reasoning layer just because it already exists in a system.

This is the gap between organisational control and data subject control.

An organisation may be able to say, “We control our data.” But the individuals, customers, employees, patients, citizens or businesses inside that data may reasonably ask, “Did we control this use of us?”

For highly sensitive data, the default should not be “connect it and see what value emerges.”

The default should be purpose, permission and proof.

Purpose: why does the AI need this data?

Permission: who authorised this use, and can affected people or organisations object?

Proof: can we audit what the AI accessed, what it produced, who received the response and whether the outcome was appropriate?

These questions do not stop innovation. They make innovation safer.

AI connected to private data could help organisations work faster, reduce administrative burden, find important information, improve decision-making and deliver better services. In areas like health, agriculture, emergency management, biosecurity, finance and regulation, the potential benefits are significant.

But those benefits depend on trust.

And trust depends on more than security claims.

It depends on transparency, consent, proportionality, access control, auditability and human accountability.

The next phase of AI will be defined by private data. That is where some of the greatest value will come from. It is also where some of the greatest risks will emerge.

So the governance question should no longer be limited to:

“Is our data being used to train the model?”

The better question is:

“Can my data be used by the model to create responses for others, and who gets to decide?”

Because once AI can reason over private data, access is no longer a technical detail.

It is the risk.

Source: Larry Ellison, “Oracle’s Vision and Strategy: Oracle AI World 2025”, YouTube,