AI Agents: Navigating Reliability and the Future of Web Interaction EO 2026-05-04

AI Agents: The Promise and the Peril of Current Technology

尽管市面上涌现了大量宣称能执行任意网络任务的 AI 代理产品,但现实往往是,它们在首次尝试时便会失败。这种不确定性和低可靠性,尤其是在需要多步骤(如10步、20步甚至50步)的复杂工作流程中,会导致错误率迅速累积,使得任务的整体成功率变得很低。演讲者对此深表忧虑,认为行业似乎已经开始容忍甚至“正常化”这种不确定性和低可靠性,尤其是在产品交付方面。作为一名产品构建者,他认为将一个宣称无所不能但实际却无法一次性正常工作的代理产品交付给用户是不可接受的。他强调,如果一个代理产品无法在第一次尝试时就可靠地工作,那么它就“不够好”。这种对“含糊不清”和“不可靠”的容忍,尤其是在智能代理产品领域,是需要被挑战的。

Original English

There's basically a hundred different

agent products out there that are saying

that like this can do anything on the

web and you try it once and it doesn't

really work. If we think of a 10step, 20

step or 50step workflow, even if the

accuracy at each step is like 90%, the

10% error rate compounds very quickly

and so the overall success rate of a

task of a workflow is like quite low. So

that's one of the reasons why the

technology is not there yet to do long

horizon workflows. It feels like we have

started normalizing and developed a

tolerance for non-determinism and low

reliability in shipping products. The

product builder in me is like really

annoyed that how is this okay for

someone to ship an agentic product where

they say that this can do anything but

you try it the first time and like it

doesn't work. I push back on that

getting normalized especially with

aentic products. If it's not good enough

to work on the first try it's not good

enough. The true differentiator is in.

Abhishek Das's Journey: From Medicine to Building AI Futures

Abhishek Das,Ytorii 的联合创始人兼联合 CEO,讲述了他如何从一个具有医生世家背景的家庭,最终投身于工程和人工智能领域。尽管来自医学世家,但他本人却对血液感到恐惧,这使得他早早地决定不追随父辈的脚步。他对科学及其严谨的科学方法(Scientific Method)——从提出假设(hypothesis),设计实验(experiments)来验证或推翻假设,到得出结论(conclusions)并提出下一轮假设——深感着迷。这种对系统性探究过程的欣赏,引导他进入工程领域。在 IIT Roorkee 的本科学习期间,他遇到了许多才华横溢的人,但很快意识到自己对电气工程(尤其是电力系统方向)的兴趣不高。这促使他在大一结束时做出了一个重要的“叛逆”决定:不再过度关注电气工程,而是将大量时间投入到学习编程(programming)和软件构建(build software)上。IIIT Roorkee 当时拥有强大的编程社群和文化,特别是 SDS Labs 这样的团队,他们热衷于为校园构建各类应用程序,这极大地激发了他的热情。从零开始构建东西,并看到用户如何与之互动,这种“多巴胺”的反馈成为了他持续前进的动力。他从小就怀有创业的愿望,尽管在本科毕业和博士毕业时都曾认真考虑过,但因各种原因未能实现。最终,他认为解决世界上的有趣问题并实现自己关心的愿景(vision),而非仅仅为他人工作,是促使他创业的关键。

Original English

My name is Abhishek Das. I'm the

co-founder and co-CEO of Ytorii. So with

Ytori, we're building agents that can

take actions and complete tasks on users

behalf on the web so that you can focus

on whatever is most meaningful to you.

The three co-founders, we're all AI

researchers by background. It is a

bigger bet than just another agent

company. So Faith Lee and Jeff and they

were excited to support us.

I come from a family of doctors and

medical practitioners. So it was a bit

of a an irony that I'm scared of blood.

And so like the choice was pretty clear

that like yeah I'm not going to pursue

medicine. That's when I ended up

deciding to to pursue engineering.

Science and the scientific process and

method really appeals to me like the

whole life cycle of coming up with

hypothesis then designing experiments to

validate or invalidate those hypothesis

drawing conclusions and then coming up

with the next set of hypothesis. I think

that is a very neat sort of process and

method. I found that really inspiring.

And then I went to IIT Riy for my

undergrad. That was a an amazing sort of

learning experience. Met some of the

smartest people I know. And so first

year of college I was getting good

grades but very quickly I realized that

like electrical engineering especially

this was more geared towards like power

systems etc was not where my interest

was and so at the end of first year that

was sort of the first major sort of

rebellious streak in me where I decided

that okay like I don't see a future in

electrical engineering I'm going to stop

paying as much attention to it. So that

was like a fairly sign significant fork

in my in my life in some sense and

instead I ended up spending a lot of my

time learning programming and how to

build software. That's when I got into

building software and applications very

seriously. What also helped was that

IIIT Rurki had even at the time this is

like almost 13 14 years back had a

really strong programming club and

culture. In particular, there was this

group called SDS Labs, which was a group

of like 10 15 coders from from every

year who were just tinkering and hack

like building a ton of applications for

the internet for the rest of the campus.

Seeing how users use it, building

something from scratch and putting it

out there and seeing how users interact

with I think that was like a dopamine

hit that kept like sort of fueling this.

And I would stay up nights to to build

this, to add features, to like improve

it and and so on, right? Especially

because I was surrounded by people who

were as obsessed with this stuff as I

was. And that was like extremely

extremely motivating.

To be honest, I wanted to start

something of my own for as long as I can

remember. I had strongly considered it

at the end of my undergrad, at the end

of my PhD, and for various reasons

didn't end up doing it. So, it was just

a matter of time. And the main reason

for it is that like there's lots of

interesting problems in the world to

solve to go after. Like I did want to

push on that vision that I care about as

opposed to working on somebody else's

vision.

Reimagining the Web: Proactive Agents and Enhanced Accessibility

Over the past few decades, the fundamental interaction model of web browsers has remained largely unchanged: opening a page, clicking, scrolling, and typing. However, there's a significant opportunity now to reimagine this experience. The future, as envisioned by Abhishek Das, involves interacting with our AI assistants (AI assistants) that can take actions and complete tasks on our behalf on the web. This means moving towards a slightly higher level of abstraction, where digital agents (digital agents) work proactively in the background for us, handling mundane "digital chores" so we can focus on more meaningful and interesting tasks. This is seen as a more achievable near-term reality than physical agents. This collaborative model aims to improve overall productivity, rather than simply replacing humans. Furthermore, this evolution will democratize web access, making it easier for individuals, such as parents who may not want to learn every new website's interface, to accomplish tasks online simply by directing their AI assistant.

Original English

Over the last two or three

decades, web browsers by and large have

stayed the same, right? Like we open a

browser, we open a web page, we click

around, scroll, type stuff, etc. There

is an opportunity now to reimagine what

that experience looks and feels like.

We're going to be talking to our AI

assistants that take actions and

complete tasks on the web. And a lot of

it is going to be agents that work in

the background in a proactive manner for

you. That's what the future looks like.

And that is how we approached it. It

felt like before physical agents become

a reality um digital agents will become

a reality like the timeline for digital

agents is shorter than for physical

agents. If you think about interacting

with the web maybe like 5 to 10 years in

the future, it is going to be at a

slightly higher level of abstraction.

Instead of us having to do digital

chores ourselves manually, it lets us

focus on tasks and stuff that's more

meaningful, that's more interesting to

us. Like if we can delegate all the

mundane stuff to AI assistants, AI

agents on our behalf, it lets us focus

on stuff that's more more interesting to

us. So it's more like humans and agents

working together to overall improve

productivity less so that like these

agents are going to like replace humans

and then humans won't have anything to

do. But part of it is also just making

it accessible to more people. Like my

parents for example no longer have to

learn every new website and how to

operate it, right? Like if they can just

tell an assistant that this is what I

want to do on this particular website

and it does it for them reliably, then

that's awesome, right? So it makes it

more accessible for more people.

The Critical Path to Trustworthy AI: Craft, Taste, and Rigorous Development

Building reliable AI agents that can handle complex, long-horizon workflows is an ongoing challenge. A key aspect is developing the ability for these agents to recognize their own mistakes and backtrack effectively to find alternative solutions. Ytorii invests significant effort into building robust evaluations (evals) and guardrails (guardrails) to monitor production queries, quickly identifying areas where agents excel and where they need improvement. Given the sheer volume of websites, it's impossible to train on every single one; thus, models are expected to make mistakes, much like humans do. The crucial differentiator lies in whether the agent can recognize, backtrack, and correct itself. This requires a meticulous approach to product development, moving beyond mere functionality to focus on craft (craft) and taste (taste) – ensuring intuitive and well-designed user experiences. The practice of dogfooding (dogfooding), where the team rigorously uses their own product internally, is vital for refining their taste and understanding what constitutes a truly "awesome" or "magical" experience. The principle of attention to detail (attention to detail) is paramount; if users can trust the visible parts of a product, they are more likely to trust the unseen mechanisms. Projects like Grad-Cam (Grad-Cam) illustrate the importance of interpretability (interpretability) and providing "proof of work" – showing not just the answer, but the steps taken to reach it. This transparency is fundamental for building user trust and delivering reliable products that don't just appear, but are intentionally and meticulously built.

Original English

In this day and age, there's basically a

100 different agent products out there

that are saying that like this can do

anything on the web and you try it once

and it doesn't really work. And there's

also this notion of that like if you

usually works right like if you try it

10 times then maybe like three times or

like five times it it does the right

thing I push back on that getting

normalized. agents are basically making

a sequence of decisions. Like if we

think of a 10step, 20 step or 50step

workflow, even if the accuracy at each

step is like 90%, the 10% error rate

compounds very quickly. And so the

overall success rate of a task of a

workflow is like quite low, right? And

so that's one of the reasons why the

technology is not there yet to do long

horizon workflows. Being able to

recognize when it makes mistakes and

backtrack from that to then uh go down a

different branch is really really

important. We put in a lot of effort

into building evals and guardrails. Like

every single production query that a

user runs goes through a fairly

comprehensive set of evals that lets us

quickly identify where these agents are

doing well versus not, which domains

need more work and so on. That's one

aspect of it. And because we're in this

space of web agents, right? Like agents

that can do actions and tasks on the

web. It will never be the case that we

will be able to train on every single

website that's out there. Like there's

new websites coming up all the time. The

number of websites that exist in the

world is already pretty large. So we

will always be training on a finite set

of websites and improving these models

there. Like people make mistakes on new

website, click on the wrong buttons,

etc. all the time, right? Like so it is

very natural to expect models to also

make mistakes. But when it makes a

mistake, is it able to recognize and

then backtrack and correct itself to do

the right thing is a fairly important

ingredient in the recipe of like how we

train and build and ship these models.

But the other part is is more

ecosystemwide where like it feels like

we have started normalizing and

developed a tolerance for

non-determinism and low reliability in

shipping products. I don't like the

normalization of slop and

non-determinism and poor reliability

especially with agentic products. Yeah,

if it's not good enough to work on the

first try, it's not good enough. We take

sort of an 80/20 approach to it. Like

there is always the prioritization

question of like okay there are 100

features that we could be building. What

are the top 10 that we need to focus on?

Like some of those are informed by users

and what what users are are telling us

what they're asking for. But very often

there are ways to build product that

users may not be asking for. But if you

built it and a lot of intuition goes

into identifying what those features

might be. Then users feel seen and they

feel like oh this is someone who is

listening to us. Even though that's not

exactly what they asked for initially.

I'll give you an example. The feature on

iOS or Android that like anytime you get

a two-factor authentication SMS, it auto

reads your SMS and fills it into

whichever app asked for it. It is hard

to imagine like a user asking for that

feature. But it saves a few seconds

multiple times a day for people all

across the world. But it's like a tiny

thing that makes users feel seen like oh

someone is actually giving thought how

to reduce these tiny paper cuts in our

in our day-to-day life. That's really

important. So like it is a marriage of

intuition with what users are actually

asking for. In a world where it's very

easy to come up with first prototypes

using these coding LMS the true

differentiator is in taste and craft in

how intuitive and welldesigned the

product is. One thing we do in the team

that helps with that I think is we take

uh dog fooding our own product very

seriously like every single week we have

an hour hour and a half docked out for

dog fooding new features in the product

at any given point of time we're running

like tens of experiments internally and

maybe like one of them will ship to the

um production version of the product

that external users will see. So like

constantly dog fooding our our own

product is a way to refine our own taste

for like okay what is good versus bad

what awesome or magical feels like. Like

anything else um a lot of reps uh to

build that muscle is like one way to go

about it.

The grad cam project was led by one of

my labmates. I was sort of a supporting

author on that paper. I was 25 when we

did that paper. It's been extremely

wellreceived. I think 20 30,000

citations is quite non-trivial at the

time. Interpretability in like around

deep learning models was like a big area

of focus still is to this day. And so it

was motivated from that that like okay

like these models especially

classification models to start with that

go from like images to classifying it in

one of thousand or 10,000 categories.

What part of the image are they looking

at to make those predictions? Right?

There is clearly some signal coming from

the image itself and then some signal

that may be coming from the label that

the classification model is predicting

and how can we combine the two develop

better intuition for what part of the

image the model is looking at. To this

day, it seems to work quite effectively

across a bunch of tasks and models. Like

with AI models, it is important for

models to be able to convey not just the

final prediction or the final answer,

but also the proof of work. like what

are the steps that went into coming up

with this final prediction or the final

answer. And so Grad Cam is like one

manifestation of that. But even in how

we build the the scouts product today,

like you can set up these scouts and

agents to monitor the web for something

and they will generate these reports and

notify you when they find something

that's of value to you. But there is a

button in the UI that lets you inspect

the work that went in in behind the

scenes like which websites were were

visited, what did the agent actually

look at to pull out this piece of

information and that gives you a glimpse

into the work that went in behind the

scenes to put this together. It is very

very important for trust building for

users to be able to trust that yes this

is a reliable product.

A lot of our time and attention in how

we're building our product at UTI goes

into thinking about how should we build

the product so that we don't make the

same mistake. Whenever we ship

something, we have to get it right. It

has to really work. It has to be

reliable. Users have to trust that it

works well. If we put attention to

detail into parts of the product that

users can see, then the user is more

likely to trust the parts of the product

that they cannot see. Right? like

everything awesome that we see around

us, it's like individuals or groups who

put in a lot of hard work and attention

to detail to build that. So I think we

should approach everything that we are

building with that kind of philosophy.

It takes time to build something

meaningful, to build something right, to

bring a vision of the future to life and

building like delight delightful and

reliable product experiences. It doesn't

just appear out of nowhere.

📌 文中提及的人物和组织

公司/组织: Ytorii, IIT Roorkee

产品/模型: Grad-Cam

关键字: ai-agents web-interaction reliability product-development ai-workflow