For my Agent Reality Lab project, I have created a git repo here so that anyone may clone it if it gets somewhere. I also searched Kaggle to find a dataset that can be used for my project and decided to use this one. The data is from anonymized behavior logs of the OTTO webshop and the app. I’m using this data to create an experiment environment for the multi-agent system I’m going to learn how to build and observe.
The very first thing to do when you are given a new dataset is always: EXPLORE! Whether your job is a Data PM, analyst, engineer (absolutely), or any other data role needs to be able to “see” what the data looks like.
So, the dataset is in JavaScript Object Notation (JSON) format, which is semi-structured. The file size is also huge! I might not be able to simply open it in a text editor. Pandas always comes in handy.
This is basically what it looks like if you open it. I used Python to look at the first three records. Each record is a unique session with multiple events in it. You can also use the JSON Pretty website to explore small files or a particular record.

Imagine that when you visit an e-commerce site, you will be given a session dedicated to you. During that session, you can click on multiple products (aid), add items to your cart, and finally place an order.
Now, I want this JSON file to be transformed into a structured dataset. How do I do that? Python also helps. It can explode and normalize JSON data format.

Now, we’re talking about the user journey and funnel. What I’m exploring next is the proportion of clicks, carts, and orders. All in order.
Because the data size is huge if I cannot say humongous, I only checked the first 50k sessions.
The file contains of
Total events: 216,716,096
Unique aid: 1,855,603
From the 50,000 sessions, I got
clicks: 2,395,745
carts: 180,595
orders: 44,770
Now, we’re talking about the marketing funnel. It is important in order to really see what can be improved for each stage in the funnel. For example:
- If we want to increase customer awareness of certain products, we may want to increase the views on the product page.
- If the customer has any desire to purchase the product, then the product will be moved to the cart.
- We need the customer to take the final action to purchase the product by placing an order and making the payment.
We can analyze the conversion from each stage of the customer journey and give recommendations on what or where to improve, best accompanied by what intervention we can do to improve.
Next, what am I going to do for the project?
The research question: Does a multi-agent architecture reproduce human behavior better than a single LLM agent?
I’m going to compare the following:
- statistical baseline
- traditional ML
- single LLM
- single agent (LLM + memory)
- multi-agent system
- multi-agent + learned behavior model
My next post will be about building the statistical baseline.