Exploring Beyond Data Lineage: Tracking What AI Agents Actually Do

J

Jaimin Suketu Patel

Guest
I have devoted quite a considerable amount of time to the concept of lineage within the context of data governance. This concept is relatively straightforward: if I am working with the data, then I should have enough understanding of its origins, its journey, and its destination.

Such approach has proven itself effective when it comes to figuring out the course of the data.

However, things get more complicated when an AI agent comes into play. An AI agent does not necessarily consume or transform data; it may use the data for decision-making or invoking other systems or tools.

In other words, I do not just need to know the destination of my data anymore. I also need to understand: How the data has been used by the AI agent?

When AI Agents Move From Data to Action​


Let’s take a simple example:

A user requests help from an AI-assisted support agent to address a billing problem. The agent looks up the user’s account details, reviews their transaction history, finds a duplicate charge, and makes an API call to process a refund to the billing system.

In terms of data lineage, I could potentially have fairly good visibility into the data around the customer and the billing. I would know its origin, how it was retrieved, and which application it went to.

However, now imagine somebody returns four months later and asks why such refund was processed. The fact that I know the origin of the customer data is not going to tell me much of the story.

I would like to know who the agent was, what data it had at hand, what it did based on this data, what made it do that, what other tools it used for that, and how exactly that was done.

From Data Lineage to Action Lineage​


This is where I started thinking about another type of lineage alongside data lineage: Action Lineage

I think of it is pretty simple. If data lineage helps me trace where the information came from and how it moved, action lineage helps me trace what happened after an AI agent used that information.

Here is how it might look like in the case of billing refund:

Tracking an AI Agent's journey from request and context to action and outcome


Here, I’m not suggesting that action lineage needs to become another large governance framework. For me, it’s simply a practical way of connecting the dots.

As long as an AI agent is able to take some action that impacts on any customer, employee, transaction, system, or business process, I would like to know how it got there; what inputs it had, what decisions it made, what controls were put in place, and what results followed.

Data lineage still matters. Action lineage simply continues the story.

The Log Doesn’t Tell the Whole Story​


For instance, if i’m investigating the refund, I come across this in the log:

Code:
13:23:04 - billing.refund(account=948321, amount=93.33)

That’s useful because I now know that the billing API was invoked and a refund was initiated. But what it doesn’t tell me is why it happened.

How did the process get initiated by the customer request? What kind of information did the agent rely on? Why did the process determine that a refund should be provided? Did the agent have the authority to provide the refund autonomously? And was the refund successful?

The answers to all these questions may perhaps exist in different systems somehow. It is the matter of connecting all these dots into one picture.

This is the difference I am concerned about. In addition to knowing whether there was a request to use an API, I would also like to know how this happened.

Governance is Also a Piece of This Journey​


There’s another part of this chain that I think is important: the controls that influence what an agent is allowed to do.

It is possible that sensitive data was detected and redacted before the action reached the agent. It is possible that there was a policy regarding whether or not the agent was able to perform a specific action. It is also possible that human approval was needed for the refund above a certain amount to be completed.

As I am attempting to understand the history of the action at a later time? Yes, which is why these are the things that also need to be considered. By finding out if the control was assessed, if the action was performed or blocked, and whether human approval was needed for the action to proceed, I get more insight into the situation.

That is equally, I believe, an important aspect of action lineage because it involves not only the tracing of the action, but also the controls that guided it.

Why It’s Even More Important as Agents Become Autonomous​


This becomes even more significant as we provide our AI agents with the capacity to take actions rather than merely producing outputs.

As soon as the agent is able to perform activities such as record updates, refund processing, ticket creation, messaging, access modification, or any other type of workflow, it doesn’t really help anyone if I tell them that the “agent did it” when they ask what happened.

The greater the level of autonomy that the agent has, the more it becomes essential for me to know the trail that led to its actions such as the information used, decisions made, controls employed, tool used, and outcome achieved.

This is where I think the combination of data lineage and action lineage comes into play. While one helps me understand the journey of the data, the other helps me understand the journey of the action.

Concluding Remarks​


Data lineage isn’t going anywhere. We still need to understand where data came from, how it moved and where it was used.

But as AI agents move beyond generating responses and start taking actions across systems, I think there’s another trail we need to be able to follow: the actions those agents leave behind.

To me, the difference is clear:

Data lineage helps me understand what happened to the data. Action lineage helps me understand what the AI agent did because of it.

As we give AI agents more ability to act, being able to trace both sides of that story will only become more important.
 

Thread statistics

Created
Jaimin Suketu Patel,
Replies
0
Views
3
Back
Top