> Identify AI model overuse with User Insights
[AUTHOR: Ayush Kumar]
[DATE: 30/09/2026 13:00]
[LANGUAGE: EN]
When we launched User Insights last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.Our latest update adds something our users have been asking for: context. Since launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.User Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.Why AI usage is hard to understandConsider a team that has routed its internal AI traffic through AI Gateway. After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.There could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.Tokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.Helping teams find where AI models are overkill The model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.That gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.The overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:Is this model appropriate for the task?Is the extra capability improving the result?Would a faster or less expensive model produce an equivalent outcome?Is the issue limited to one workflow, user, or agent?From there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.These insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.The Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.Understand what people are using AI forTask analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.This provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.The answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.Understand the full cost of a taskSome tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.The first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.Turn insights into auto routingOnce a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.For example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.In addition to our updates to User Insights, the Auto Router is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.Instead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.To learn more about the Auto Router and sign up for the closed beta, read the blog post here.The Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.How User Insights classifies trafficEach conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.The categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.The Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.The pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.The classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.The tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.The flow looks like this:Connect usage to users, teams, and toolsTask categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.This works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.For custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.For custom applications, the request metadata might look like this:The request body contains the model and messages for the conversation.With Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.Get started with AI Gateway User InsightsAI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.User Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.Learn more with the AI Gateway User Insights documentation . Open AI Gateway in the Cloudflare dashboard, and use what you learn to make more targeted model and routing decisions.