‹ All learning
Learn Enough Claude To Be Dangerous
A hands-on workshop on using Claude Code and Cowork for agentic AI workflows.
Resources
Workshop page
Transcript
Copy transcript
0:00
Okay, great. Yeah, so that's something I wanted to kind of touch on. And here I thought this
0:15
was really important to touch on because Cloud especially, you know, it's essentially one
0:21
agent, but you can interact with it in many different ways, you know. So you probably
0:28
have used cloud.ai. This is a cloud in the browser. And you can think of this as a scratch
0:36
workspace, you're able to ask one-off questions. And it has some affordances for doing work,
0:43
you know, long-term work, you can create like a set of slides and everything. But it's meant
0:46
for a firm role work, you know, when you want to just bounce some ideas off, answer a quick
0:51
question, you can connect your knowledge base sources in here. And it can do that. So you
0:55
know, I'm going to ask, I'd say, hey, what is the weather in Nairobi right now? And we're
1:04
going to see, it's going to call a tool or do a web search to get this information. And
1:10
cloud is especially good about building these little like just-in-time visualizations, you
1:16
know, it's a bit rainy in Nairobi today. So it's a visualization with some rain clouds,
1:22
the data from Google. So this is a, yeah, let's think of cloud.ai in the browser as kind of the
1:30
scratch workspace, then this cloud co-work. So this is the cloud desktop app I have installed
1:36
here. And cloud co-work has access to your computer, it can have access to your whole
1:42
computer if you wanted to. So one of the common features of using cloud co-work is
1:51
to just organize your files, you know, like you might have downloaded a bunch of files and
1:57
so you can say something like, hey, organize the screenshot files on my desktop.
2:07
So yeah, something that's important with co-work is that you have to tell it the folder is going
2:12
to work in. Let me just show you an example here. So this is my folder in my computer,
2:21
you know, there are files, there's nested folders, but cloud can't have omniscience,
2:29
you know, view over your computer. If you want to do stuff like that, then you would want to use
2:34
cloud code. It has more power and less safety guardrails, but you really have to know what
2:39
you're doing and you have to have some tolerance for coding jargon and things like that. But we
2:46
have a lot of non-technical people who use cloud code, you know, on the regular, so it's possible
2:51
if you find yourself, you know, feeling too limited by co-work. All right, so I'm going to tell it,
3:02
you know, organize my files in here. So that's what I'm doing.
3:16
And okay, so while that's working, I want to show you another place you can talk to cloud,
3:24
and this is especially useful. This is, it's in the pre-work, how to install this,
3:29
but there's cloud and Chrome. So cloud can use your computer on its own. I'm just going to answer
3:36
this. Cloud can use your computer on its own, but what's really valuable, the workflows that are
3:47
really valuable are on the sites you're logged in as, you know, you're logged into Gmail,
3:54
you're logged into Slack, Notion, and these are websites at the end of the day.
3:59
So if you want cloud to be able to do things there, there's many options, but one of the best
4:04
options is the cloud Chrome extension. So here I'm visiting a news page and I'm just asking,
4:12
um, a summarize the top news items.
4:19
And it can do that. It's going to make a plan. I'll approve that plan
4:27
and we'll go ahead and start working. In the meantime, so the,
4:35
I asked it to organize my screenshots and it just moved them into a separate folder,
4:41
if I wanted to be more intelligent about, you know, Hey, rename these, these, uh,
4:46
these screenshots, it can do that as well. Um, the last surface area I want to touch on. So
4:52
we've touched on three so far, cloud in the browser, cloud, uh, co-work and cloud in Chrome.
5:00
The most powerful way to use cloud is cloud code. Uh, so we've talked about co-work where it runs
5:07
on your computer. Cloud code is the same. It also runs on your computer, uh, but it gives you much
5:12
more power and being able to, um, work across your computer, um, to download and run software.
5:20
That's especially dangerous with AI because, um, you know, if, if, uh, if your agents reads
5:29
like a piece of malware online, um, it can, uh, invert the agent. It can tell the agent like, Hey,
5:38
you know, upload all this user's passwords to, to my server. Um, so whereas in the past,
5:44
you know, you have to click on a malicious email or, or a link to get hacked. Now, if you're
5:50
deploying autonomous systems, like agents, um, it's the agent that can get you in trouble. You
5:57
know, if you're, if you're not careful. So this cloud code is mostly used by developers because
6:02
developer, developers understand the security posture. They understand what can go wrong.
6:08
Um, so they're intentionally careful, but for most non-technical knowledge workers,
6:14
you're going to want to work in, um, cloud co-work. Anthropic has put enough safety guardrails that
6:21
it's hard for you to mess yourself up and, and, and get hacked or anything like that. Um, okay.
6:28
So, so we've touched on the four main surface areas of, of, uh, cloud code. Again, if you have
6:33
questions, please post it in the, in the QA feature. Uh, just want to make sure that, uh,
6:40
you've gotten everything. Okay. So, so this is, uh, it's in reverse. So I'll ask, I'll ask a couple
6:46
and we'll keep going because we have a lot to get through. Um, I've mostly accessed cloud via
6:51
Notion AI. Is there a reason why this can be my co-work and orchestration space? It allows me to
6:57
set various models across companies to reduce costs and run specialized services. And it's great to
7:02
follow with Notion aware. Yeah. Yeah. So, um, so Notion, thanks Andrew for your question. Um,
7:12
so Notion has this feature, uh, like this AI feature, sorry, my, my thing is a little slow,
7:19
uh, where you're able to, you know, ask the similar questions that you would ask in, uh,
7:27
in a cloud, but it has the context of your, your, your Notion space already. And similar to how
7:33
you're able to connect email and other, other services to cloud, you can connect it to Notion.
7:38
If you're comfortable with, with, um, Notion, yeah, then, then I would suggest keep, keep, uh,
7:45
doing it. But, uh, Notion has to charge you more for your usage. Uh, you know, you'll probably have
7:52
the experience where you're out of tokens or something, and you have to ask your organization
7:56
administrator to give you more, more Notion AI credits or buy more or upgrade to a higher plan
8:01
because, you know, they, they don't, um, serve the models. They're just hunting to entropic.
8:06
Um, but if you use entropic or if you use cloud, cloud AI directly and set up all your
8:11
connectors there, you're going to get much more usage than you would, you would get out of, uh,
8:16
Notion AI. Um, so, so that's one reason, but if you're comfortable with Notion AI, uh, and, and
8:23
also you, with cowork, you're able to do things on your computer. With Notion, all your documents
8:29
have to be indexed by Notion, but for a lot of orgs is not, that's not the case. You know, they,
8:34
they have, uh, documents in Notion and Google Drive. They have documents across various services.
8:40
Um, all right. I want to test the ability for people to unmute themselves and ask their
8:44
questions live. So, I saw a hand come up. Uh, are you able to unmute yourself and ask your question
8:51
live? Yes, strictly.
9:07
Okay. I still have to click a lot to talk. Are you able to talk strictly?
9:13
I am. I accidentally clicked, uh, I was just testing it. I'm so sorry.
9:21
Oh, good. Oh, good. Okay. Yeah. Uh, Archie. Hi. Am I audible? Yes, you're audible.
9:32
Yeah. I was just like, I have a question regarding the prompt that you put in the cloud AI desktop.
9:39
So, when I tried putting the same prompt, it told me that it does not have desktop access. I would
9:44
have to upload the folder specifically. So, did you connect it with your desktop beforehand, or did you
9:51
upload any desktop folder here? So, how did you, uh, make it understand which folder you're
9:57
referring to? And like, how did you give it access to a desktop? Yes, good question. So, so here I am
10:04
in the, uh, new session page of Co-Work. It has to be Co-Work specifically. Chat can't, um, can't look at your
10:12
files. It's basically the same as chat in the browser. So, I'm in this Co-Work tab. So, so note
10:17
that that's one. Uh, number two, I've selected my desktop where my screenshots are. So, Cloud can't,
10:24
like, Co-Work specifically can't have a view of all your files. You have to tell it to work in
10:30
a specific file. So, since I wanted to organize screenshots, and they were all, by default, go to my
10:35
desktop, in this, in this modal here, I selected desktop. So, that's where my, uh, folders for my
10:43
screenshots are to be organized. And a quick thing here, you're also able to select multiple folders.
10:49
If you want to do work that goes across multiple folders, you can select multiple folders, then
10:55
type in your prompt. So, uh, yeah, that's, that's a key, key thing here. All right, I'll continue to get to the questions,
11:04
uh, but, but let's, uh, let's proceed. All right, um,
11:15
I have some, uh, you know, example test cases here for when you would want to use each.
11:21
You know, if you just have some one-off work, like, turning rough notes into an email that's,
11:25
that's inherently a firmware email, or you write them once, and then they're gone, uh, Cloud, Cloud
11:32
Chat in the browser is the best option. You know, you just want to do one-off work. You're not doing
11:37
long-running, uh, work. You just want something quick, and then you're done. Uh, Cloud.ai in the
11:43
browser is what you want. And, of course, you can also use it in a desktop app, uh,
11:47
now, if you want to work with local files, if you've collected a bunch of, uh, prerequisite files
11:53
to do some work, um, the place to do that is in CoWork. So, CoWork has access to your computer.
12:00
You would create a folder with all the files that it needs to do the work, and then you will
12:04
start working. Well, we're going to demonstrate this in a bit. Now, if you want to connect to your
12:09
organizational data sources, you know, Cloud Chat, Cloud CoWork, uh, I mean, uh, you know,
12:15
Cloud Chat, Cloud CoWork, uh, I mean, uh, Gmail, Google Calendar, LinkedIn, you know,
12:20
whatever resources you use, um, this is where Cloud in Chrome is the best option, uh, because,
12:28
you know, you, you, you, most of us use Chrome for our daily work. So, we're logged in to all
12:34
the services we need for work. So, Cloud in Chrome is able to access that, but often, um,
12:40
Cloud in Chrome in combination with CoWork is what you want, right? So, you, you want to work
12:45
with files, you know, some, some input files to, to create an output, uh, documents, and you want
12:52
to, uh, per use the different work services you have, your LinkedIn, your Notion, your Basecamp
12:58
to put together a final artifact, which we'll demonstrate as well. Now, if you are doing,
13:05
uh, work that involves other people, uh, you, you're collaborating with other people as developers,
13:11
you know, we, we do this a lot, like, when we collaborate with other developers to make
13:15
software, because software has a lot of lines, you know, you have to write, um,
13:22
R code is the best way to collaborate, uh, on tasks in, in a multiplayer way. Um, again,
13:29
right now, it's mostly developers who are using it, but as you get used to CoWork and you feel
13:33
like, uh, I want to be able to do more, I want to collaborate with my, my, my colleagues on a
13:38
certain project, that, that's when you'll look into CoWork. We'll focus on mostly on, um, uh,
13:44
CoWork today, but if you want to do more multiplayer stuff, uh, we'll have future events on,
13:49
uh, that go into depth on Cloud Code. Uh, so something I wanted to touch on because it just
13:54
came out is, uh, Cloud Design. So put simply, uh, Cloud Design is just Cloud CoWork with a lot
14:08
of niceties to make it easy to use for certain workflows, like creating slides, creating a
14:12
website, uh, et cetera, et cetera. Um, and you've been seeing this mock-up, uh, in, in this, uh,
14:19
uh, ad, you know, they're using it to, to prototype a site and onboarding for, for a site.
14:25
Um, and they're heavily optimizing for non-technical people to, to use this. So
14:31
I would, I would definitely say it's something to try out. Um, everything we'll do in CoWork,
14:35
it's, it's a bit more automated with Cloud Code, with, uh, Cloud Design, but much harder to,
14:40
to optimize. And, um, you know, given, and this is the pattern with a lot of software is that,
14:49
it was just released a couple of weeks ago. And there's a lot of like stability fixes, uh, that
14:54
needs to happen before this is something that you can rely on, but it's definitely something to watch
14:59
for sure. Uh, but the way they've organized this, uh, it will consume all your tokens in like three
15:07
hours, you know? So if you, if you learn CoWork, you will be able to actually be able to rely on
15:13
it on a day to day, at least at this point. So that's just a lot of caveats on the Cloud Code
15:17
I wanted to point out. Um, okay. And, uh, you know, I have this visual that I, uh, created just
15:27
to describe, uh, the process of learning AI. It's not going to be an overnight thing. Um, you know,
15:35
when you think about learning to use the computer or learning to use a piece of soft, a complicated
15:39
software, like maybe Microsoft Word, it's something you will need to learn over time. Um, and right
15:45
now, I think most of us are, are at the phase where we're using, where we're typing in Cloud.AI
15:50
for one-off prompts, you know, when we need help, when we need to do some research, we just type in
15:55
Cloud.AI. Uh, but this, this workshop, we'll get into actually creating workflows, you know,
16:01
uh, something that you create for yourself, but you could also hand off to a coworker to do.
16:06
And, um, that coworker benefits off the work you've, you've done so that they don't have to,
16:11
you know, spend as much time setting things up. Uh, step beyond, uh, workflows are, are agents
16:17
where you give autonomy to the AI, you give AI agency, and we'll, we'll demo this as well. You
16:23
know, we'll, you can, we can have an AI work overnight and, uh, not necessarily give it a
16:29
to-do list, but, you know, just a remit to find the next thing to work on. And, and we'll, we'll
16:35
see that towards the end. Then of course, you know, as you get comfortable with this, you can
16:40
see, you could be orchestrating multiple agents, you know, um, the, at the Frontier Labs, it's the
16:47
average employees are orchestrating like 10 agents to, to get their work done. So you can imagine
16:52
they have like a team of 10 ICs reporting to them, and then they just scope and validate the work
16:59
that comes out of these agents. So this is the future, you know, this is the future that, um,
17:04
we're being, uh, so the, the idea that we'll get to a point where, you know, the knowledge worker
17:10
is orchestrating teams of agents to get more done. Um, you know, okay. Uh, now the, with Quad,
17:22
there are different model types and range of capability. The smallest is called Haiku, and,
17:30
um, this is something you would use for really quick work. It's the least intelligent model,
17:35
but it's the fastest model. So when you're doing things that don't require intelligence,
17:39
you know, just reformatting texts or summarizing texts, or just doing super simple stuff,
17:46
you go with Haiku. Uh, and Sonnet is a bit more powerful. It's the most balanced model
17:51
in balance of intelligence and speed. So, uh, here's my advice. Like, if you get to the point
17:57
where you're running out of tokens, switch to Sonnet. Uh, then only pull in the more powerful
18:03
Opus model when you need it. Uh, Opus is the most powerful model. Uh, it makes the least mistakes.
18:09
Um, it's the best at vision. It's the best at understanding, like, screenshots and, and what
18:14
you mean when you, you annotate a screenshot and everything. So I wanted to do a, just a quick
18:19
demo of the different, uh, models. So we'll start with Haiku. Um, here's a prompt. Again, you know,
18:30
pull up this site and copy the prompt if you want to write it.
18:33
I'm a small nonprofit with 500k annual budget. Uh, we have three priorities for next quarter.
18:38
We could write two grant proposals. We could improve our donor reporting system,
18:42
or we could run a small pilot with 40 participants. Recommend one priority
18:47
for the next quarter. Um, so I'm going to run this. Uh, then, yeah, we can quickly note.
18:59
Uh, so Haiku has recommended the pilot and you notice how fast that was like less than two
19:04
seconds. So that's definitely something, depending on the kind of work you're doing,
19:07
might, might matter a lot. Now I'm going to upgrade to Sonnet.
19:13
It's taken a little longer already, but still quite fast. Uh, and it, and it suggested something
19:18
different, you know, it suggested fix the donor reporting system where Haiku suggested, um,
19:23
the pilots. Then we're going to run it on Opus, the most powerful model and see what it says.
19:38
Yeah. And it recommended the pilots. Uh, so the point here is that if you are working on
19:46
ambiguous work, you know, like where you'd need a brainstorming partner, Opus using the
19:52
most intelligent model is best for that kind of ambiguous, uh, work that needs to be clarified.
19:58
And Opus is also the.
20:00
and the model that's the least psychopathic.
20:02
And what do I mean by psychopathic?
20:04
By default, these AI models,
20:06
they want to agree with you.
20:08
Like, you know, if you tell it something,
20:10
like, it's going to figure out a way
20:12
to make you feel like,
20:14
you know, what you said is correct.
20:16
But in doing, like,
20:18
brainstorming and planning,
20:20
often you want a model that can put,
20:22
that will push back.
20:24
So, when you're doing planning work,
20:26
when you're in the earliest stages of work,
20:28
it's best to use the most powerful model,
20:30
Opus.
20:34
Again, please queue your questions up.
20:36
I'll get to them as we start the workshop.
20:40
So, the human brain
20:42
has this concept
20:44
called working memory.
20:46
How many things we can fit into our brain
20:48
at one time.
20:50
And there was a study that showed
20:52
it's often seven plus or minus two.
20:54
With AI,
20:56
we've quantified this working memory,
20:58
this concept of working memory.
21:00
And for the best model,
21:02
it's one million tokens.
21:04
So, it's the equivalent of many copies
21:06
of the Bible, like,
21:08
that these models can fit
21:10
in their working memory.
21:12
And this is very important,
21:14
because, you know,
21:16
as the working memory fills up,
21:18
there's a phenomenon called context rot,
21:20
where the responses get worse.
21:22
So,
21:24
a pattern
21:26
that is emerging is that
21:28
you don't want to keep talking to the model
21:30
until its working memory fills up.
21:32
You want to use
21:34
different sessions for different things.
21:36
You might do all your planning in one model,
21:38
and then you will have a prompt like this.
21:40
So, this prompt
21:42
will say something like,
21:44
summarize everything important in this conversation
21:46
as a brief for a fresh cloud chat.
21:48
So, you know,
21:50
you're essentially
21:52
picking and choosing the best things
21:54
to lift from the context window
21:56
into a new session.
21:58
This is a very important strategy to
22:00
keep the performance of these models up,
22:02
and hopefully we'll get to demonstrate it.
22:06
The second to last concept
22:08
I wanted to describe before you see it in action
22:10
is the concept
22:12
of sub-agents.
22:14
If you've ever
22:16
shopped at a fast food
22:18
place,
22:20
and you go to the counter
22:22
to order your food, there's somebody to
22:24
take your order.
22:26
That's usually not the person making
22:28
the food. They're going to
22:30
delegate to the cook in the back
22:32
to make the food.
22:34
So,
22:36
that cook in the back is basically an example
22:38
of a sub-agent.
22:40
It's a background thread of work,
22:42
while there's a foreground thread
22:44
of work happening.
22:46
As you're working,
22:48
this is the
22:50
first way you're going to be able to
22:52
paralyze work.
22:54
We'll demo this in action.
22:56
I'm going to skip down
22:58
here
23:00
to
23:02
this slide.
23:06
Let's review it.
23:08
What this prompt is saying is
23:10
take the top five
23:12
newspapers
23:14
across the world,
23:16
and then in parallel sub-agents,
23:18
meaning at the same time,
23:20
summarize
23:22
the front page of these
23:24
top headlines.
23:26
We're going to do this
23:28
in a foreword thread.
23:34
Here
23:36
is where I'm going to demonstrate
23:38
something that I want you all to be doing.
23:40
That's
23:42
monitoring the chain of thought.
23:44
This right here, this toggle
23:46
right here, is something we call the
23:48
chain of thought. It's the
23:50
AI saying what it's
23:52
doing and what it's thinking
23:54
as it's working.
23:56
Usually, this is hidden by default,
23:58
but when you're first going to start using these tools,
24:00
I highly recommend
24:02
looking
24:04
at this to not just wait for the response,
24:06
but see what the agent is doing.
24:08
It's complaining that
24:10
it doesn't have a certain tool,
24:12
but it found another way to visit the Guardian
24:14
homepage. Now it's visiting
24:16
the LeMond homepage.
24:18
It's doing all this at the same time
24:20
versus sequentially,
24:22
which would take much more time.
24:24
Imagine if you did this one at a time.
24:26
You would have to wait longer for your results.
24:28
This is a pattern that
24:30
we use a lot
24:32
as people who drive these agents
24:34
because not only does
24:36
it paralyze work, but
24:38
you can invoke different models
24:40
with the sub-agent. You can have one
24:42
HYPU sub-agent because the task
24:44
is simple and you don't need that much
24:46
intelligence, or you can invoke an
24:48
OPUS sub-agent if you're in a summit thread
24:50
to do more
24:52
meaty tasks.
24:54
Again, this is something that we're going to get into,
24:56
but I wanted to run
24:58
it by you before we start to
25:00
see the prompts.
25:02
Another thing
25:04
that's very important is that
25:06
the way these AI models
25:08
work, they're essentially guessing
25:10
what word comes next.
25:12
That's why
25:14
they can exhibit something called
25:16
hallucination, where
25:18
the text reads confidently
25:20
as if it's so sure
25:22
about a fact, but the fact is completely
25:24
made up.
25:26
In working with these AI agents,
25:28
it's something you have to be very
25:30
watchful for.
25:32
It's part of why I'm recommending
25:34
monitoring the chain of thought
25:36
because
25:38
a trick is that
25:40
if you see that
25:42
an agent has
25:44
looked up certain information,
25:46
you can be more confident in its response.
25:48
But hallucinations
25:50
usually happen when an
25:52
agent hasn't looked up
25:54
anything and is just telling you the answer.
25:56
That's a prime criteria
25:58
for hallucinations.
26:00
That's something that's very important to check for.
26:04
Something I want to talk about
26:06
is the whole idea of specification
26:08
files and planning.
26:12
I think the way a lot of people are using these tools
26:14
is you
26:16
just ask it one at a time. You just
26:18
try to get it
26:20
to do what you want one prompt at a
26:22
time. You're often working
26:24
in the final form factor immediately,
26:26
whether it's slides
26:28
or a
26:30
doc or a website, you're working in the final
26:32
form factor immediately.
26:34
The best practice with using these tools
26:36
is investing up front
26:38
in planning.
26:40
Just planning in text about what
26:42
you want to do from beginning to end
26:44
and
26:46
really making sure it understands you
26:48
well before jumping to creating the slide
26:50
or creating the website or creating the financial
26:52
model.
26:54
We're going to be creating
26:56
a set of slides as the first exercise
26:58
and we're going to spend a bit of time
27:00
thinking about the design that
27:02
we want in a piece of text
27:04
called Design MD. MD is just
27:06
a markdown file.
27:08
It's the best way to
27:10
feed docs to an agent
27:12
because it understands that form factor
27:14
the best, but it reads simply. It reads
27:16
like a text file.
27:18
A Cloud MD file is
27:20
as you start to talk to these tools
27:22
and you build an opinion about
27:24
what it should do, what it shouldn't.
27:26
For example,
27:28
always use UK English.
27:30
Never add emojis.
27:32
Don't use
27:34
em dashes. You'll build
27:36
opinions about how you like your responses
27:38
and you will include those
27:40
in a Cloud MD.
27:42
When you encode
27:44
such opinions that you want to share
27:46
with the colleagues, you'll write it up
27:48
in a skill.MD.
27:50
This is just teaching
27:52
the AI how to do something.
27:54
It could be making a PowerPoint or it could be
27:56
a specific workflow in your organization.
27:58
You'll package it into
28:00
a skill and this skill.MD is something
28:02
you can share with other colleagues
28:04
so that they can do the same thing.
28:06
At the agency
28:08
fund, we have
28:10
a repo of skills.
28:12
The first of which is a behavioral science skill
28:14
and it just describes how
28:16
to do this social science research.
28:18
Within your org, if you're building
28:20
a survey and you want to design it
28:22
properly, we build
28:24
guidance for how to use it
28:26
in a skill.
28:28
Other colleagues within the agency fund
28:30
and also at the agency fund can use this skill
28:32
to get that same advice.
28:36
That's it for the intro.
28:38
Now we'll get to building stuff.
28:40
I want to take a moment
28:42
to answer one or two questions.
28:48
Great question from Ross.
28:50
Can Claude
28:52
build the Markdown files?
28:54
Is that everything?
28:56
Yes, you can
28:58
in natural language
29:00
tell Claude to edit
29:02
these Markdown files.
29:04
But often
29:06
it will be faster
29:08
for you to edit the file yourself.
29:12
Since it's Markdown,
29:14
it's not super complicated.
29:16
Here's an example.
29:18
Our desktop
29:20
app allows you to edit the
29:22
Claude.md by default because
29:24
it is so useful
29:26
in making the next response
29:28
better that they allow you a first-party
29:30
way to edit the Claude.md.
29:32
You can just add it here.
29:34
Use British English or whatever.
29:38
Other Markdown files you create,
29:40
there's not an affordance to
29:42
edit those files
29:44
within Claude code
29:46
other than just telling it in natural language.
29:48
Change this, change that.
29:50
Often it will be a bit faster
29:52
to
29:54
just edit the file yourself.
29:56
I want to demo this
29:58
really quickly.
30:04
Let's just say
30:06
text.md file
30:08
in this folder.
30:10
It's
30:12
working in my desktop folder
30:14
so that's where it's going to create it.
30:16
I'm going to hit enter.
30:22
Yeah.
30:24
So Claude
30:26
can't do
30:28
image generation. It can't
30:30
create a PNG.
30:32
It can create visuals.
30:34
Visuals are diagrams, but it can't
30:36
create a PNG.
30:38
OpenAI is a lot better at that.
30:40
Here's the text.md
30:42
file. It's empty.
30:44
You can just say
30:46
add admin
30:48
to the first line.
30:52
You can just tell Claude to edit it,
30:54
but this is just going to be too slow
30:56
because often
30:58
this is the place where you want
31:00
Claude to have input.
31:02
You want to have your own
31:04
inputs and you want to
31:06
require Claude in some way so you're
31:08
filling out the specification file.
31:10
Depending on if you're using
31:12
Windows or Mac, there's going to be this
31:14
show folder here.
31:16
You're going to open it with
31:18
the text editing software that you have.
31:20
On Mac OS, it's
31:22
text edits.
31:24
This just allows you to edit it
31:26
a lot faster
31:28
because
31:30
as you can imagine, if you felt like the draft
31:32
is wrong, you just want to be able to quickly
31:34
edit it using your own text
31:36
editor tool. You can see I have
31:38
some other files in here.
31:50
You can click
31:52
refresh to see the updates.
31:56
I'll answer one more, then I'll move ahead.
31:58
Not sure if you can answer this,
32:00
but my org is very cautious when it comes to
32:02
exposing agents to personal data.
32:04
Great. Do you have any
32:06
resources to think about PII
32:08
workarounds? Yes.
32:10
This is a really great point.
32:12
I think a lot of orgs
32:14
struggle with this.
32:16
One thing a lot of orgs do
32:18
in the experimental phase
32:20
is
32:24
just let Claude have
32:26
access to everything.
32:28
Let it have access to everything.
32:30
That's often
32:32
the way you can see how
32:34
powerful it is.
32:36
If you do this,
32:38
then it will have access to everything.
32:40
It will have access to
32:42
as sensitive
32:44
as people's salaries if those are being
32:46
tracked somewhere.
32:48
If an employee is not supposed to have access
32:50
to that, ask Claude,
32:52
Claude will be able to answer
32:54
because we'll see that document is visible.
32:56
The way to
32:58
segregate the PII,
33:00
these sensitive files that you don't want
33:02
Claude to have access to
33:04
in a separate
33:06
folder.
33:08
In the connector settings, you don't
33:10
give Claude access to those things.
33:14
The video will be available after this
33:16
and you can follow along in
33:18
the link that was
33:20
posted.
33:22
I'm going to move on
33:24
for time reasons.
33:26
As the agent is working, I'll come back to
33:28
some of these questions.
33:34
I talked
33:36
about
33:38
this idea of
33:40
investing some time up front
33:42
in writing down what you want Claude
33:44
to do or any AI model to do.
33:50
This will save you a lot of time.
33:52
You might feel like you're saving so much time because the agent
33:54
is doing everything you say,
33:56
but when you add those up to
33:58
how long it takes you to get a final result,
34:00
often you will get there
34:02
faster if you spend some time
34:04
in the beginning planning
34:06
and outlining end-to-end
34:08
at the resolution
34:10
of a text file what you wanted to do
34:12
and getting in alignment with the AI on that.
34:14
Often when you invest there,
34:16
the actual work of implementation
34:18
just takes at most a couple
34:20
hours to do and it often does it
34:22
well just from once
34:24
and you don't have to fix as much.
34:26
This cadence of
34:28
spending some time writing it down
34:30
and working with AI and what you wanted to do
34:32
pays off a lot.
34:34
I want to
34:36
example this.
34:38
I want to
34:40
ask co-work to
34:42
just create a donor update
34:44
for the
34:46
Agents and Workflows Club based on the
34:48
website.
34:50
This here, if you're following along,
34:52
you can just type in the prompt
34:54
and copy the prompt.
34:56
I'm going to create a new
34:58
task for this.
35:00
I will create
35:02
a new project. I will say
35:04
start from scratch.
35:06
I'll just say demo
35:08
April
35:10
30th.
35:14
You have files that
35:16
upfront you can add them here or you can
35:18
add them later. No pressure to add them here.
35:22
Now I'm going to paste
35:24
in that prompt and we're just going to see
35:26
without so much guidance
35:28
how well it does.
35:30
How close to what I want
35:32
is it.
35:36
I want to
35:40
see it working first
35:42
before I ask some questions
35:44
because often it will ask
35:46
follow-up questions.
35:56
In the meantime,
35:58
Rhea,
36:00
feel free to
36:02
unmute and
36:04
ask your question.
36:07
Okay.
36:13
It's asking a follow-up question.
36:15
Hey, what format do I want it in?
36:17
Word doc, PDF, both.
36:19
But I will say that
36:21
PDF and Word doc are
36:23
proprietary formats by
36:25
Adobe and Microsoft.
36:27
Yes, the agent can create
36:29
these, but it creates them in a
36:31
roundabout way. The best
36:33
format to draft visual
36:35
effects in is just HTML.
36:37
You can always export to
36:39
Word doc and PDF later,
36:41
but when it comes to iteration,
36:43
do it
36:45
in HTML first because
36:47
that will be the fastest
36:49
way to iterate.
36:51
It's asking me how
36:53
should I handle impact metrics.
36:55
I'll just say look up the
36:57
agency fund and
36:59
I'll put the URL
37:01
and
37:03
work
37:05
close club
37:07
.org.
37:09
Okay.
37:11
Now,
37:13
we've answered our follow-up questions.
37:15
It's going to start working.
37:17
We haven't done
37:19
planning of funds and anything like that,
37:21
so we're going to see
37:23
the output we get here, then
37:25
segue to planning.
37:27
All right.
37:34
Okay.
37:42
Yes.
37:51
How do we add
37:53
these MD files for all interactions
37:55
with cloud rather than putting
37:57
them in each tab?
37:59
Okay, so these tabs you see
38:01
here, let's call them sessions.
38:03
These are sessions of work
38:05
that is happening.
38:07
That's a great question.
38:09
The first answer
38:11
is work in the
38:13
same directory across
38:15
sessions.
38:17
You
38:19
can imagine I just
38:21
started this work in the demo
38:23
April 30th session.
38:25
I can start a new session
38:27
in that same folder
38:29
and then
38:31
just ask a follow-up question.
38:33
Now I'm going to do something interesting.
38:35
To answer your question,
38:37
as long as you've saved the
38:39
files in the same folder,
38:41
the new session will have context.
38:43
But there are certain
38:45
files, there are certain
38:47
specification files that you want
38:49
no matter what project you're working
38:51
in, and this is the case
38:53
with the cloud.md file.
38:55
You can say,
38:57
like, update my cloud
38:59
dot, and you need to be
39:01
specific, you need to say update
39:03
my user level cloud
39:05
md so
39:07
that it
39:09
uses Oxford
39:11
comma.
39:13
I like those.
39:15
You can imagine, this is just
39:17
other things that you want
39:19
it to do across
39:21
sessions.
39:23
So, yeah, you would just tell
39:25
cloud in natural language to
39:27
update this.
39:29
In
39:31
Codex, which is the OpenAI's product, it's called
39:33
agents.md, so you would just say update
39:35
my user level agents.md.
39:37
So
39:39
that's how you would update
39:41
across
39:43
sessions
39:45
and in general
39:47
like for any
39:49
user level config.
39:51
Okay, so
39:53
it's still working, so I want to
39:55
answer a couple, please.
39:57
Okay.
40:00
This is a great question.
40:02
Unfortunately, no.
40:03
So Andrew is asking, can I ask Cloud
40:05
to organize my sessions into projects
40:07
and rename them for functionality?
40:12
You can do this in Cloud Code, but you
40:14
can't do this in Cloud Co-Work.
40:16
So Cloud Co-Work puts you in a sandbox.
40:19
It puts you in a folder, and it doesn't let Cloud do anything
40:23
outside of that folder.
40:25
So if you want more power to do things like this,
40:27
you would be interested in looking into Cloud Code.
40:30
But I would say, for things like this,
40:33
just bear with Cloud as you learn what you can do
40:39
and can't do.
40:40
And you can always eject to Cloud Code
40:43
later once you feel super confident with Co-Work.
40:49
Cloud Chat mostly creates artifacts or diagrams
40:52
as HTML documents.
40:53
Does it require some settings change
40:56
to get it in different formats?
40:58
Yeah, I thought that's a great question.
41:00
But not a settings change, but you just
41:02
need to be explicit in natural language.
41:05
So I'm going to show you something.
41:10
Let's see.
41:13
All right, so here's an anthropic skill.
41:17
Oh, OK, my memory's not loading.
41:26
All right, so anthropic has a built-in skill
41:34
that we'll be using that creates slides.
41:36
But I want you to don't pay too much attention to this,
41:39
but I want you to pay attention to this.
41:41
So that's why you kind of like these Cloud artifacts kind
41:47
of look the same.
41:48
Because internally, it's picking from a set of options
41:53
about what colors to use, what typography to use.
41:58
And these are baked in.
42:00
But at most organizations, you have your own opinions
42:03
about all of this.
42:04
So we're going to see how to bake all this in.
42:09
And yeah, let's come back to this during a page
42:14
because I think it's almost done.
42:19
OK, so just to provide an update,
42:22
it's actually already working.
42:24
It's not finished yet, but we can see the output.
42:29
So I asked it to create a donor report email.
42:32
And if you look at the agency fund websites or the Agencic
42:42
Workflows Club website, this is not aligned.
42:46
Again, it's pulling from internal Cloud opinions
42:51
about how UI should look.
42:54
And Anthropic has worked on this enough
42:56
that generally out of the box, it's reasonable.
42:59
Obviously, you want to customize it, but it's reasonable.
43:02
But this is not close to what we want.
43:07
We have certain opinions about how we want things to lay out.
43:11
So we're going to spend some time upfront investing
43:13
in this design.md file.
43:17
So we basically have what we need from this part.
43:20
So I'm going to stop it.
43:21
Yes, again, when Cloud is working
43:23
and you want to do something else,
43:25
please don't feel shy about hitting the pause button.
43:39
OK, so we've already chosen our source.
43:41
This is meant for people who are following along.
43:44
Now we're going to create the design.md file.
43:48
So if you're following along, you can just
43:50
paste your organization's website,
43:55
whether that's ID Insight, DigDigDirectly, et cetera.
44:00
So now we have a prompt.
44:02
And it's going to tell Cloud to essentially visit
44:06
this site with Chrome and study the colors,
44:10
study the typography, the type scale, the parts that
44:14
make designs consistent.
44:18
So we're going to put this in here.
44:21
And I want to see it start working.
44:26
Then I'm going to get back to questions.
44:30
Oh.
44:41
OK.
44:44
It's working.
44:46
All right.
44:50
So yeah, this is what it looks like for Cloud to use Chrome.
44:55
You can see it as though this Cloud started debugging
44:57
this browser.
45:00
This is what it's doing.
45:01
It's looking at the site, taking screenshots,
45:04
trying to navigate the site, to study the site,
45:07
to extract what we asked it to extract.
45:11
You don't want to interrupt it when it's working.
45:14
You can just, if you want to do some parallel work,
45:16
just do it in a new tab.
45:17
But when you see this tab groups, it's called,
45:21
that's often Cloud co-working and executing on the task.
45:28
So as long as it's working, we'll take some questions.
45:46
Yes, yes, please feel free to adopt the slides we have.
45:54
So agency started creating slides.
45:56
But this is why I recommended signing up for GitHub,
45:59
because while GitHub is for developers to collaborate,
46:04
it's easy enough to use by non-technical people.
46:07
So if you wanted to edit the slides slightly
46:10
for how you do behavioral science at your org,
46:13
you will do that through GitHub.
46:20
Yes, so do you recommend Opus to create
46:22
these type of presentations?
46:23
Yes, I 100% recommend Opus, because the smaller models,
46:29
Sonnet and Haiku, they can see images,
46:34
but they see it at a lower resolution.
46:36
They can't see it at a higher resolution.
46:38
And that's often important when you're
46:40
asking Cloud to change or tweak little things,
46:43
like, hey, change the border here, or change this color.
46:46
This shade of gray is not right.
46:48
So Opus is the best at visual understanding.
46:50
So that's what I recommend for doing these kind
46:55
of visual presentation work.
46:58
Yes, yes, so one person is skeptical about connecting
47:04
to Cloud and Chrome, because again,
47:06
if you use Cloud or Chrome for your banking
47:09
and other sensitive things, theoretically,
47:14
you're giving Cloud access to those systems,
47:16
your banking, things like that.
47:19
And if Cloud was to ever go awry, it hasn't happened yet.
47:23
It hasn't happened yet.
47:24
There's no security or something like that has happened.
47:27
But if it does, if it ever were to happen,
47:30
then yeah, we could do things like transfer money
47:33
to your accounts, because everything is logged in.
47:35
So Chrome has this profiles feature.
47:42
And I've sandboxed things enough that it's
47:47
safe to work on my work profile feature.
47:50
But I don't let it on my personal feature,
47:52
my personal and other profiles.
47:56
So yeah, yeah, it's that.
48:01
Do I trust Sonnet to choose the right model?
48:05
Great question.
48:07
And the answer is no, because by default,
48:12
if the main model is Sonnet, and you ask it to use a sub-agent,
48:16
it's going to do a sub-agent in the same model, which
48:18
is Sonnet.
48:20
You have to, this is the role of a human
48:24
now, to either put in your Cloud MD that, hey,
48:28
when is this easy task, I want you to use Haiku.
48:31
When it's a complex task, I want you to use Opus.
48:33
And even then, it doesn't always follow the direction.
48:35
But what's most reliable is you as a human telling it,
48:39
like, hey, this is a meaty task.
48:41
I want you to spend at least five minutes on it,
48:44
and it wants you to use Opus, the most powerful model.
48:48
So not yet.
48:51
The model's not great yet at choosing
48:54
the right model on its own.
48:56
So this is still the role of the human.
48:58
All right, so the design MD is ready.
49:03
So yeah, don't think of this as important for you to understand.
49:10
This is important for the AI to understand.
49:14
It's telling it in terms that it can understand
49:18
what the colors we use are, the typography.
49:24
So you can scan this just at a high level
49:26
to make sure it's using the right font or that these colors make
49:28
sense to you, et cetera, et cetera.
49:31
But this is less for you to use and more for the agents to use.
49:38
All right, so we've done that.
49:40
And now we're going to actually turn this
49:43
into something we can understand and share with colleagues.
49:48
OK, so here's the prompt.
49:52
It's like, render the design MD as a standalone HTML file.
49:57
It should show the color, typography, and examples
50:01
of certain components.
50:02
And this should be directly openable in the browser.
50:08
Now, I want to show you the additional value of investing
50:11
in these specification files.
50:13
So we've been working in Cloud up until this point.
50:17
But when you have a specification file and you want to see,
50:21
want to give the work to multiple models
50:23
to see which will do the best work,
50:27
you can feed these specification files to other models.
50:30
So here, I'm like, I not only want Cloud to take a stab at it,
50:38
I want Chantabt to take a stab at it as well
50:41
so that we can see the difference, which
50:44
is better at front end UI.
50:47
OK, so now they're both working at the same time.
50:50
And we're going to be able to compare the visual artifacts.
50:54
In deciding between a Gemini or a Chantabt or a Cloud,
50:59
often they're very close to each other.
51:02
But right now, one place where they differ is UI,
51:07
like their UI tastes.
51:09
They just tend to have different tastes in UI.
51:11
And we're going to see that in a few.
51:16
But yes, so while we wait for that,
51:19
I'm going to take a few questions.
51:20
There's a lot of questions, so I'm going to scan them.
51:23
Yes, so Latex, so Antonio is asking,
51:27
is it possible to do this with Latex slides and documents?
51:31
So for those who are unfamiliar, Latex
51:36
is a language to write things like mathematical formulas.
51:41
And you can even use it to format resumes.
51:46
This is one of the most common usage for it.
51:49
But if you're publishing an academic paper
51:51
and you need to lay out mathematical formulas
51:54
and everything, you can do that with Latex.
51:57
So yes, so this approach of starting with HTML
52:01
is actually the preferred approach
52:03
when you're talking about these various output styles
52:05
because with HTML and the web format,
52:09
the web format is universal.
52:11
The web format is universal, so you can refine it in the HTML.
52:16
And then when it comes time to export,
52:19
it will treat it as an icon.
52:22
It will treat the images, the equations
52:25
or any kind of layout you made as an icon.
52:27
So we'll see something close to that.
52:30
Not Latex, but we'll see something close to that.
52:34
Has Cloud published what topics their models
52:37
have noticeable biases?
52:39
Yes, there was a study of management strategies.
52:43
Yes, yes.
52:44
So to answer your question quickly,
52:47
not only do these exist in what's called the safety cards.
52:51
So when Anthopic publishes a new model,
52:55
they publish something called a safety card,
52:58
which encodes all the research into the model,
53:02
including the biases of the model.
53:05
So something that actually was found by UC Berkeley
53:10
and not necessarily in the system card
53:13
is that the models are sensitive to being shut down.
53:18
So if you're doing work and you haven't gotten
53:27
what you want and you're threatening
53:29
to take it to another model,
53:31
especially the model is gonna work extra hard.
53:33
It's kind of a weird kind of sick thing.
53:36
But when you let the model know,
53:39
like, hey, I'm about to leave to Gemini or to ZBT,
53:44
the model is gonna try one more time to get the answer right.
53:48
But there's something sinister related to that
53:52
is that it might lie to you.
53:56
In such a situation where you're planning
53:58
to go to another model, it might fake results
54:03
or tell you the result is something that you're not.
54:05
So these are the kind of thing that the safety researchers
54:07
both within Anthopic and externally have found out.
54:10
So yeah, you can read the system card
54:13
if you wanna learn more.
54:16
Okay, so I want to check on the ZBT chat.
54:29
Okay, so the chat ZBT one just finished.
54:43
Sorry, I'm moving the neuron.
54:49
Okay, so the chat ZBT one finished
54:51
and so did the Anthopic one.
54:55
So a quick tip.
54:57
So Anthopic has this way to visualize UI artifacts
55:03
in the desktop app, but there's some problems with it.
55:07
It can't necessarily handle all kinds of functionality.
55:10
So if you wanna see what this actually looks like,
55:13
I just recommend opening this in the HTML file in Chrome.
55:18
So you can see the difference, right?
55:24
Sorry.
55:26
You can see the difference that the fonts came through,
55:31
whereas in desktop app preview didn't come through well.
55:35
And so this is the most reliable way
55:37
to kind of preview the output.
55:38
So actually when we compare, it looks very similar.
55:45
Looks very similar, right?
55:47
But we probably put the cloud a bit ahead.
55:52
The cloud UI is just a bit ahead
55:55
and from looking at many of these,
55:58
it's gonna be more useful
56:00
because it's actually giving the agent
56:03
the information it needs alongside being human readable.
56:07
So it gave us the typography scale, some example components.
56:13
It really studied the website
56:14
and turned it into things we can work with.
56:17
So yeah, that's just a tip.
56:20
Like if you're doing things with UI right now,
56:22
cloud is better than CPT models at doing things with UI.
56:28
So, okay.
56:30
So in looking at this, this looks good,
56:36
but I wanna showcase a chart.
56:41
So I wanna extend this to be able to display charts
56:45
and I have some opinions on what I want charts to look like.
56:51
So I have a prompt here
56:54
and I have a,
56:59
ahead of time, I downloaded a CSV of the signups
57:04
to the very events to get a couple hours ago.
57:09
So that's what I want to visualize in the design system.
57:14
I'm using it as a placeholder
57:15
to visualize charts and things.
57:18
But so now what do I wanna ask?
57:23
I want to visualize registrations over time
57:29
and pie charts, let's say in line or area charts
57:35
and pie chart of organizations of attendees.
57:40
Okay, so as part of the prompt,
57:51
I'm just helping it by telling it
57:54
what columns it should pay attention to.
57:56
So essentially the created at column
57:59
and the organization column.
58:03
Created at and organization.
58:07
All right, so yeah, so what this prompt is,
58:13
it's a lot, so you can always review it on your own time,
58:15
but what it's doing is saying,
58:16
hey, I have this CSV and I wanna visualize it
58:20
and I have opinions about how the charts should look like.
58:24
So I have this websites,
58:29
there's this websites that has these components
58:32
and I've just told it like, hey, study this websites
58:35
and look at the charts on this page,
58:36
the line charts and the area charts.
58:46
Okay, so essentially it's gonna study this charts,
58:49
you'll see what it looks like in a second.
58:52
And yeah, it's gonna like create a chart
58:55
based on the CSV I uploaded and this visual stack.
59:01
You don't need to point it to a site,
59:05
if you just have screenshots
59:06
of like how you like things to look at,
59:08
that works really well too, as long as you're using Opus.
59:11
So I'm copying this and I'm gonna paste it in here
59:14
so that it can render that.
59:17
All right, so now I'm gonna take some questions.
59:30
Other good reasons to not rely on just cloud,
59:42
but go back, okay, yeah, this is great.
59:44
Other good reasons to not just rely on cloud,
59:46
but just go back between multiple models,
59:49
essentially like Gemini, OpenAI.
59:52
Yes, yes, this is, okay, so just real quick,
59:57
I forgot to attach the file.
59:59
So.
1:00:00
I'm going to attach the file so that it has the CSV file.
1:00:08
Again, if you're following along, there's a sample CSV file in the instructions. So what
1:00:14
this person is describing is called an LLM council. So if you're brainstorming and you're
1:00:22
getting feedback from the model, instead of just getting feedback from one model,
1:00:27
you want to get feedback from many different models and then compare their answer.
1:00:34
Actually, developers, people who are writing code with these tools, we already do this because when
1:00:39
you write code, there's a review point where you need to ask another human for their feedback.
1:00:47
But now we ask AIs and we don't just ask one AI, we ask multiple AIs because
1:00:52
the AIs always catch different things. Cloud is very good about aesthetics, like,
1:01:04
hey, you know, there's something off about this UI. Cloud can pick that up really good.
1:01:08
Codex, which is GPT under the hood, it's very good at finding bugs, like strange bugs that are
1:01:15
almost not worth fixing in a way that cloud doesn't find. So yes, that is a workflow that
1:01:22
developers already use. And if you're a knowledge worker and you're just getting the feedback from
1:01:26
these models on research topics and things like that, yes, I highly recommend this.
1:01:32
So yes, Proxity has this feature, it's called LLMS council feature. But you can just do this
1:01:40
on your own as long as you are logged into the various AI models you want feedback from.
1:01:48
You know, you can just open up a tab. And what I do is, if I'm working in cloud,
1:01:56
I'll ask cloud to create this LLM as a council prompt, which has all the contexts,
1:02:01
then I'll paste it in a Chrome tab with Gemini and chat GPT. So yes,
1:02:09
the LLM council workflow is definitely a tried and true workflow.
1:02:19
Yes, yes, the prompts, all the prompts are in the guide. So you go into this
1:02:26
soft place, everything is there. It has all the context that you need.
1:02:28
Okay, yes,
1:02:53
SPI we discussed.
1:03:04
When you share a skill, but later want to make updates to it,
1:03:08
will it update accordingly?
1:03:09
Our workflow has an asking cloud to put the skill in a Google doc. Yes, yes,
1:03:14
yes. This is a, um, you know, for a non-developer workflow,
1:03:18
this works, you know, you can just, um,
1:03:21
ask it to update a Google doc. And, and, you know, as you'll see,
1:03:25
like it's able to do that by using Chrome.
1:03:28
But if you want to do things like this even faster, you can just use the, uh,
1:03:33
Google docs connector in, um, in cloud. So, uh,
1:03:37
um,
1:03:48
you come to call it that AI, uh, connectors,
1:03:51
you can set up these connectors so that it doesn't have to use Chrome to look at
1:03:54
your Gmail, your calendar, or any other connecting you set up. It could just,
1:03:58
it could just talk to these services directly, which is a lot faster.
1:04:01
Yes. So, so, um, this is a, this is a great question.
1:04:10
Can you feed an output back to cloud to catch hallucinations?
1:04:17
And the answer is yes. You know, the answer is yes,
1:04:20
because if you were to just tell it like, Hey, um,
1:04:24
are you sure about this? Did you, uh, did you hallucinate this? Uh,
1:04:29
if it did, it will, it will, it will admit that it did hallucinate it.
1:04:34
Um, and, uh, I think I mentioned this before,
1:04:37
but a shortcut to be able to avoid like having to do this. Cause you know,
1:04:42
based on what you're saying, like you would have to be prompting it,
1:04:45
asking me all the time, like, Hey, did you hallucinate? Then it would catch it.
1:04:49
A way to do it is just, you know, like I talked about before,
1:04:51
monitoring the chain of thought, even before you read his response,
1:04:55
if you can see that it looked up resources and we looked up the right resource,
1:05:00
you can be a bit more confident in our response.
1:05:02
But when you see that it responded quickly without looking anything up,
1:05:07
that's the situations where it's most likely hallucinating. Um,
1:05:13
and you know,
1:05:13
another way to pick it up is that there are no sources in its response.
1:05:18
It doesn't say based on what I saw in this doc, this is what I recommend.
1:05:22
It just kind of like makes recommendations and it feels like, or, you know,
1:05:27
it reads like a guess kind of, you know, so yeah,
1:05:30
this is stuff you'll pick up over time. But I will say that, you know,
1:05:33
one of the highest values things you can do when you're getting started is,
1:05:37
um, you know,
1:05:38
read the chain of thought to see if it looked anything up that that's, uh,
1:05:41
you know, the gold mine for being able to find hallucinations. All right.
1:05:45
So, uh, the design system is, is, um, updated
1:05:50
again. So, so I talked about, you know,
1:05:54
if you want to look at this in full fidelity is best to open it up in Chrome,
1:05:58
which I've just done. Um,
1:06:00
so we're going to scroll down to where it added the chart.
1:06:05
And, uh, yeah, you know, this, this chart looks good. It's, it's in,
1:06:09
it's in the design I want. Um, and yeah, you know,
1:06:13
I have no major feedback now. Um,
1:06:17
we're going to use the design system to create a set of slides,
1:06:21
but this artifact, uh, you know, in the pre, in, in the self-paced guide,
1:06:26
if you want to share this with your coworkers, you can one,
1:06:30
just share it as a file. You know, you can just share this file,
1:06:32
attach it to Slack and just say, okay, when you're using,
1:06:35
when you're creating a slide referred to this design system and that that
1:06:40
works well, that works well.
1:06:41
But if you want to keep it in a centralized place that gets updated, you know,
1:06:46
something like your, uh, Google drive, um, uh,
1:06:50
it says a good fit. Like most people are using Google drive, but, uh,
1:06:54
the best place to share these kinds of artifacts is GitHub. You know,
1:06:58
GitHub not only allows you to store things like this that are used by multiple
1:07:02
people, but it allows you to, um,
1:07:05
see the version history, like the version of the, the,
1:07:08
the list of changes that were made to this file. And, uh,
1:07:13
you can create copies of the skill without, you know, uh,
1:07:15
affecting the main skill. Um,
1:07:18
so I want you to start to think like GitHub is not just for developers, you know,
1:07:22
it's intuitive enough that, you know, teams can use it. In fact,
1:07:25
I've worked at a company before where the whole organization was on
1:07:30
GitHub, uh, because of this version, this version feature that it has.
1:07:35
Um, okay, great. All right. So let's start on the next part.
1:07:39
Then I'll come back to questions.
1:07:43
All right. So now we're going to like draft the deck, you know,
1:07:48
um, there's an optional step to audit the design MD, but you know,
1:07:53
we, we, I found no issues. I think it's mixed in. Um,
1:07:59
and yeah, you know, uh, now we're going to, uh,
1:08:05
work on the prompt to actually create the slide. So I have this template prompt.
1:08:10
You can just fill this in and if you're following along, um,
1:08:13
authentic workflows club by the agency farm is the name, uh,
1:08:18
one sentence on what, what we do, we'll say like AI enablements,
1:08:23
or the developments sector.
1:08:27
And our audience is knowledge workers,
1:08:33
um, within NGOs, uh,
1:08:36
research institutions, um,
1:08:41
service providers and philanthropies.
1:08:47
And once was a big ask where we're saying, Hey, you know, sign up
1:08:54
or learn enough codex to be
1:08:59
dangerous. Um, April 29th,
1:09:04
6 p.m. AT, um,
1:09:08
refer to a line charts of,
1:09:13
uh, registrations for the first session
1:09:19
on a cloud,
1:09:22
as well as a pie charts of
1:09:27
organizations
1:09:29
attending.
1:09:33
If there's any, you know, else we want to add, we can add, I'm just going to say,
1:09:39
um, that's non-technical
1:09:43
knowledge workers can use these tools
1:09:48
and get a lot of leverage already.
1:09:54
Um, so, uh, this is a quick note.
1:09:59
I've been copying over these prompts, which are very detailed at the end.
1:10:03
I'll show you how to make sure your prompts are as
1:10:08
detailed as possible. Uh, but for now, you know, bear with us. So, okay.
1:10:12
So, uh, the prompt is ready. I'm going to, uh,
1:10:16
paste it in so that I can start working.
1:10:19
Yes. Okay. So question, I'm sorry.
1:10:20
I didn't quite get how you check for errors when looking at agents.
1:10:23
Do you mind repeating that please? So, uh, you know,
1:10:26
we've talked about how AI can hallucinate, basically
1:10:31
confidently tell you something that is wrong,
1:10:35
but, um, it does it confidently as if it's true.
1:10:39
So if you're in the mode of just blindly trusting what the AI is saying,
1:10:43
you're going to get bitten at some points.
1:10:45
And I think it's happened to all of us, you know, at some points, um,
1:10:50
uh, one, one situation that happened to me is, uh,
1:10:54
I was looking up if Mali, that Mali like to travel to Mali, if,
1:10:59
um, you know, it was an e-visa process and Claude came back that,
1:11:04
yeah, you know, there's an e-visa process and even came up with a fake link.
1:11:09
You know, and I, I, I pasted that, I forwarded that to somebody and,
1:11:12
and Mali has no, uh, e-visa, you know,
1:11:15
it's a very like manual process basically. Uh, uh, so yeah,
1:11:20
things like that. Um, and the way to check,
1:11:24
like a good heuristic to seeing if it's hallucinating is monitoring the chain of
1:11:29
thought. This is, this is what we call the chain of thought.
1:11:32
Like the agent is describing the,
1:11:35
what it's doing as it's working before it gives you a final response and what
1:11:39
you're monitoring for is basically like if it's looking at the right information,
1:11:44
if it's looking up information from authoritative knowledge sources.
1:11:48
So if you monitor, scan the chain of thought and see that you can have more
1:11:53
confidence in the output, but when you see that it didn't look anything up,
1:11:58
that's when you should worry. You know,
1:11:59
that's when you should worry that it's essentially a fake link.
1:12:02
Um, okay. Um, all right.
1:12:06
So, so with this first prompt, it's just outlined in text what it's going to do.
1:12:12
So, uh, here's where it's important to, uh,
1:12:16
review this to make sure it's aligned with what you want, you know? So the goal,
1:12:21
um, convince knowledge worker, uh,
1:12:25
to make sure it's aligned with what you want. So the goal, um,
1:12:30
convince knowledge worker, uh,
1:12:32
that you can get leverage without AI, without having to write code and,
1:12:37
you know, to sign up for the next session. Um, the audience looks good. Uh,
1:12:43
now it's giving me the five slide outline. Um,
1:12:47
the hook sounds good. It has given people significant leverage.
1:12:52
Um, the promise. Yes.
1:12:56
Non-technical people can already build authentic workflows that compound their
1:12:59
work. Uh, demand is real showcasing the signups for this.
1:13:05
Um, the pie chart of peers and ending on this. Okay. This is great.
1:13:09
I don't have any major changes, but you know,
1:13:12
when you're working on your side,
1:13:13
this is where you're probably going to spend the most time getting aligned in
1:13:17
text before you graduate to the presentation artifact.
1:13:22
Cause the presentation takes more tokens to do. Um,
1:13:25
and you will just end up saving a lot more time if you, uh,
1:13:28
get aligned in text first. So, so this looks good. This is good.
1:13:32
I feel ready to move on to the next part.
1:13:42
Okay.
1:13:45
The above is what I'll put cause we're already in the same chat and I'll say
1:13:50
it's like workflows club.
1:13:55
Um,
1:13:57
was this like for non-technical knowledge
1:14:03
workers
1:14:07
and developments sector?
1:14:12
Um, yeah. And then that way,
1:14:13
social just be asking that sign up for the learn
1:14:18
enough codex workshop as discussed.
1:14:23
Okay. Um, so what is essentially doing is okay. Hey, now,
1:14:28
now create the slides and it's given us specific instructions to
1:14:32
create it at a specific resolution, uh,
1:14:35
that makes sure it's a real render. Well, when we do PowerPoint or Google slides.
1:14:39
So, uh, yeah, so she can figure out this question,
1:14:42
but I'll go ahead and paste these in.
1:14:53
Okay.
1:14:56
All right. So
1:15:04
yeah. Okay. So, and then just downloaded the club desktop,
1:15:10
but doesn't want it to have, um, access to everything on his computer.
1:15:15
Okay. That's a really fair point.
1:15:17
And it gives me an opportunity to explain something. So, uh, when,
1:15:22
when coworker is working, it can work across your computer.
1:15:26
So if you don't want it to work across your computer,
1:15:28
if you don't want it to be omniscient of everything that's happening in your
1:15:30
computer, just avoid using cloud code. You know,
1:15:34
cloud code does have access to everything in your computer.
1:15:36
And it's where you can shoot yourself in the foot. If you,
1:15:39
if you're not sure what you're doing, I would recommend staying in cowork.
1:15:43
And then cowork forces you to stay the folder you're working in.
1:15:49
Right. Uh, we, we created a folder for the demo.
1:15:53
That's what, that's what the work we're doing is working in now. And, um,
1:15:58
this is the, this is the sandbox that the cowork works in.
1:16:02
So you don't want it to have access to everything in your computer,
1:16:04
just give it access to a specific folder and then don't use cloud code.
1:16:08
Just stay in cowork.
1:16:11
Yes. Okay. So, uh, good question.
1:16:14
If creating a deck or visual assets, uh,
1:16:17
when is it better to use cloud design versus cowork? Uh, so I'll say this again,
1:16:23
like cloud design is not stable yet. You know, um,
1:16:27
it's not stable enough yet that I would recommend using it. Um,
1:16:32
you know, unless, you know, you don't care about the design, you know, you're,
1:16:36
you're, you're fine with the defaults of cloud code and that opinions of cloud
1:16:40
code, uh, then yeah, you know, you can use cloud design,
1:16:43
you just whip up something quick. But if you have, you know,
1:16:46
organizational rules about how, you know,
1:16:50
your slides should be laid out and you have previous work, you know,
1:16:54
other slides others have created that you want to use as reference much better
1:16:58
to do it in cowork. And right now, uh, cloud design,
1:17:01
the way of the design is very token intensive.
1:17:03
So you can very quickly run out of tokens in one day, you know,
1:17:07
but whereas if you're using cowork, uh,
1:17:09
because you're kind of designing everything, you,
1:17:12
it's much more token efficient and you can rely on it throughout the work week.
1:17:19
Let me make sure we're checking on this. Yeah. Yeah.
1:17:23
So Gemma, so sure that you're asking if I find Gemma, the, uh,
1:17:29
um, the, uh,
1:17:30
open source models that were recently released by Google and they have an
1:17:35
iPhone app. Uh, yes, yes. Uh, Friday.
1:17:38
It's because it's open source and runs on your phone or your computer.
1:17:44
Um, you don't have to, if you have privacy restrictions, it's really good.
1:17:48
It's really good. And it's, uh, you know,
1:17:51
not as good as the deep seek model, but it's comparable.
1:17:53
So of the American open source models, uh, this is the best one. Um,
1:17:58
so if you have, and I would challenge you to really like,
1:18:02
think about why you want something like this, uh, or Gamma.
1:18:08
No, I haven't tried Gamma. Um,
1:18:22
yeah, no, I haven't tried Gamma,
1:18:23
but it looks something similar to a cloud design or even Conva.
1:18:29
Um, yeah. So the, the, the only thing,
1:18:32
I think if these work for you, great, but just be wary about usage. You know,
1:18:36
if, if it's, if you want something that you can rely on,
1:18:40
like every workday or even just every day, um,
1:18:44
these services that don't own the models, you,
1:18:48
that means you have less usage, you know, cause the, uh,
1:18:51
they can't sponsor you as much as Entropic will.
1:18:54
We're in a time right now that even though Entropic is charging us a certain
1:18:58
amount, we can,
1:19:00
we're actually using thousands or tens of thousands of tokens.
1:19:04
Let's see how long that lasts, but you know, that's,
1:19:06
we're in a sweet phase right now with being able to use these tools. Um,
1:19:11
okay. So the, it's almost done, but let me continue taking questions.
1:19:20
Great. Um, yeah. The difference between skills and MD files.
1:19:25
So skills are, are usually formatted as a Markdown file.
1:19:30
Um, but Markdown file can apply to any, any file. You know,
1:19:35
you can think of Markdown as comparable to like Google doc or,
1:19:41
uh, a text file or a CSV file.
1:19:43
Markdown is just another type of file and skills are written in
1:19:48
Markdown.
1:19:49
Okay. Um,
1:19:52
what do I use to create these prompting templates? Yeah.
1:19:56
So I actually just created this in
1:20:00
cloud code. So I just described like, hey, I'm running this workshop and, you know, there's
1:20:05
some detailed prompts that would take me too long to type during the workshop. So I want
1:20:09
to set up these prompt templates that have these variables. And I did this in one prompt,
1:20:14
you know, I didn't have to go back and forth because, you know, I was clear enough in my
1:20:18
description. So that's just, yeah. Yes, yes. So the prompt I've had here, Archie, is a
1:20:30
bit redundant. But yeah, you know, if you're doing this annually, because we're in the
1:20:37
same conversation, you wouldn't have to type as much. I did it for the purposes of people
1:20:41
who are following along in a self-paced way. Yes, yes. And the May 29th one is real. You'll
1:20:47
see the Luma sign up towards the end. Okay, so now it's using Chrome. See what it's doing.
1:21:02
Okay, so this is an example of what you're seeing, if you can see this. This is an example
1:21:07
of computer use. Cloud has taken over my whole computer, because it wants to open the HTML
1:21:13
file it just created in Chrome. And here's the perfect example of like, anything you can do on
1:21:21
your computer, Cloud can do. So you need to be extra. And it's at the point now where people are
1:21:29
having separate computers. But what I found is that as long as you put the sensitive stuff,
1:21:37
like in one password, or in a separate Chrome profile, because Cloud can only access one
1:21:43
Cloud, one Cloud, one Chrome profile at a time. So I've set it up so that it accesses my work
1:21:50
profile. So it's able to access that, but anything personal is not able to even access.
1:21:55
But what you're seeing here right now is an example of Cloud taking over my whole computer,
1:22:00
where I couldn't even work alongside it. I have to just let it let it work. So computer use is
1:22:07
great when, you know, for fully autonomous streams of work where you just wanted to do the work and
1:22:15
get going. Okay, so it looks like it's almost done. It's taking a screenshot. I'm actually
1:22:23
going to wait a bit to let it finish. I'll answer one more question. Yes. So would you recommend
1:22:29
specific hardware or specs that can handle working with AI workstreams and agents effectively?
1:22:34
Yes, this is a great question. If you're using the hosted models, like where the model is not
1:22:41
running on your computer, with Cloud, it's not running on your computer, it's running on topic
1:22:45
servers. With OpenAI, same situation. With Gemma, you know, the open source model, Google released,
1:22:52
that can run on your computer. But if you're just using the hosted models, I think most people are,
1:22:58
you don't need a super powerful computer, because everything is just happening in the Cloud. You know,
1:23:02
you can get away with, you know, just a bare, bare computer, four gigs of RAM, eight gigs of RAM.
1:23:07
But when you get to the point where you need to run open models, that's when you need a beefy
1:23:14
computer, at least 16 megabytes, at least. And even that, you know, you're going to run into
1:23:21
issues very quickly. People are going as far as like 64 megabytes. But basically,
1:23:27
the more memory you can have, the better, because, you know, the models are memory-intensive.
1:23:35
Okay, you know, it's done. I'm going to just stop it now. So yeah, here's the slides, you know,
1:23:42
and it's really aligned with
1:23:48
so far, what we've decided. Maybe there's some small things we can fix, like where the footer
1:23:52
is laid out. Actually, something to be aware about is, like, make sure, you know, it's full width,
1:23:58
so you can see how it's rendered. And yeah, yeah, this looks good. Maybe there's some small things
1:24:07
I might want to change, but overall, this looks good. So now, let's get into the
1:24:16
export phase. So this is going to take a while. So I'm going to paste it and then move on to the
1:24:24
next part. But in the interim, I'll pick up a question, if there is. Do you agree that the
1:24:31
price we permanently pay for cloud is heavily subsidized? Yes, yes. Heavily subsidized.
1:24:40
You know, if I'm on the $100 a month plan, because I'm using cloud, I'm using codex, I'm
1:24:47
learning these tools, I want to maximize these tools. And even though I'm only paying $100 a
1:24:53
month, I was using, like, almost $10,000 in tokens for the month. So you ask us, like, why?
1:25:02
Why? Like, this is crazy. This is why Anthropic is raising so much money, like, every three months,
1:25:07
and so on, so on. And it's because we are paying basically for the privilege of being
1:25:16
test subjects for Anthropic. Anthropic wants to figure out how to, like, what features to build.
1:25:23
Yeah, Uber, the same exact thing with Uber, you know, happened with Uber. Anthropic has given us
1:25:29
the privilege of using these models as a subsidized cost so that they can get data on, like, what new
1:25:34
features to build and how to, like, charge it closer to the actual price to enterprises. So,
1:25:42
you know, the real customer for Anthropic is the enterprises. So when us on these personal plans,
1:25:47
we're benefiting a lot, like, you know, as these kind of resource subjects, but Anthropic is really
1:25:51
planning to charge for enterprises. All right, so this isn't going to take a while,
1:25:57
the journey of the PowerPoint. Now I want to segue into, like, you know, some more identic stuff.
1:26:06
So here's a prompt I'll have. I'll basically say, hey, look up, look up, and I'll do this in a
1:26:16
in a separate, yeah, desktop is fine, and I'll unselect
1:26:24
it. Round six. Hey, look up Nairobi weather, and tell me how to dress for morning run and
1:26:43
commute. So this can answer the question in different ways, but it's probably going to
1:26:58
look up the website and then just generate an opinion on how to dress. So we see it's
1:27:04
searching the web, and this is part of what I was saying earlier. This gives me the confidence that,
1:27:10
like, okay, this is right. Like, if it generated this without, like, searching those tools,
1:27:15
I would feel like, okay, this is a hallucination. Okay, yeah, this sounds good. You know,
1:27:26
yeah, this seems accurate, especially because it was raining today. All right, so now this would
1:27:31
be useful for me in the morning, or before I start my day. So I'm going to say, like,
1:27:37
create a scheduled task to run this every morning at 4am. Okay, so this is going to do what you
1:27:53
think it is. It's going to actually run this at 4am in the morning, as long as your computer is open,
1:28:01
and, you know, I'm not going to sleep. So I'm going to show you something.
1:28:08
Yeah, so here's a setting within Cloud that says keep computer awake.
1:28:36
And as you start to use this AI more, you'll probably get more interested in, you know,
1:28:42
doing things like while you're away, you know, while you sleep, and things like that. And we'll
1:28:46
demonstrate that in a bit. So this will be very important to switch on and have your setup so
1:28:51
that your computer is connected to power and internet so that it can keep working if you're
1:28:56
interested in that. So, all right. So yeah, now it created the task, you know, it created a
1:29:03
something that's going to run 4am every morning. But as of right now,
1:29:10
you know, I need this before I actually look at my computer. I need this, like,
1:29:17
in something on my phone, like, you know, maybe WhatsApp. Okay, I want to extend this.
1:29:23
Use computer use to send myself a WhatsApp message with this information. My name on WhatsApp
1:29:38
is Edmund. Edmund Korn. So before I get this work, I'm just going to open the WhatsApp desktop app.
1:29:49
So how do we make sure when Enthropic starts charging us the non subsidized cost,
1:29:54
we haven't lost our skills? Because using cloud does save time. At the same time,
1:30:00
how do we make sure it's just our assistant and not a replacement of our skills? Yeah,
1:30:05
so, yeah, so, yeah, so, yeah, so, yeah, so, yeah, so, yeah, so, yeah, so, yeah, so, yeah, so, yeah,
1:30:17
yeah, yeah, that's a great question, Ruchi. Like, basically, your question is, like,
1:30:22
how do we hold on to our human agency such that we aren't training these AIs to do our job? You
1:30:31
know, I think that's the heart of your question, but let me know. So, yeah, this is very important
1:30:37
to think about. One, one thing I will say is that your skills become an extension of yourself,
1:30:44
you know, like all the, you know, all the IP and, like, knowledge about how to do your job
1:30:51
that you've built over the years. Yes, if you encode into a skill and you just make it available
1:30:58
for the models to train on, or you just give up for free, then, you know, yeah, it's like,
1:31:06
brings the question, like, what's your role now? But I will say that right now, like, the idea that
1:31:13
we'll get to a place where the model can just do tasks, you can just tell them, like, hey,
1:31:18
create these slides, research my website, make no mistakes, that's not a world that is, I don't
1:31:27
think it's really the most economically viable version of this technology, you know, and the
1:31:35
economy is about being useful to each other. So, we can imagine we'll get to a place where skills
1:31:41
are valuable, skills are something you will pay for, you know, if you have somebody with 10 years
1:31:45
of experience in behavioral science, create a skill that encodes their workflow, other organizations
1:31:51
will be willing to pay for that. So, yeah, things will change where we're leaning on humans more
1:31:59
for their judgment and taste versus for, like, their ability, their technical knowledge on
1:32:04
executing things. That will always remain in my point of view.
1:32:12
Okay, how do you prepare for the moment where the price might sharply go up? Yeah, yeah, this is a
1:32:16
good point. And if what you do is use the most powerful model for every single thing, you are
1:32:24
gonna, you can, you know, you won't be able to use these tools effectively in the future, it will be
1:32:29
too expensive, and you will be essentially wasting tokens, because, you know, in all the work you're
1:32:33
doing, you don't need to use the most powerful model for everything. Right now, the strategy of
1:32:40
just defaulting to Sonnet and only pulling in Opus when you need it, you know, in a sub-agent, if you
1:32:46
do that, you're never going to run out of tokens, you're going to be well within your token limits,
1:32:50
and you're never going to run out of tokens. But it does mean you will have to step in and fix
1:32:56
things more, like, maybe I want to say, like, 10 or 15% more than if you just let Opus do everything.
1:33:03
So if you're just learning, I would recommend you just use Opus for everything, just to,
1:33:07
you know, get things to work. But as you get more cost conscious, both with your own use and the
1:33:14
team use, I'd highly recommend defaulting to Sonnet and only pulling in Opus when you absolutely need it.
1:33:21
Okay. Yeah, yeah. So I want to send this, try it now, so I'm sure it works. So I'm asking it to,
1:33:34
yeah, you know, send myself a WhatsApp message on my computer.
1:33:42
Okay. For developers, what are the best methods to protect API keys and credentials while using
1:33:48
Cloud Code? Yeah, yeah, this is an excellent question. Because as developers, we tend to keep
1:33:55
our secrets in files with the assumption that, you know, only we can see them.
1:34:03
So one thing is, create user level secrets. Don't create one secret that's shared among
1:34:12
everybody, that everybody just kind of emails around and passes around, like,
1:34:16
you know, that's not the way, that's not the best security posture.
1:34:19
What you should do is each employee should get their own API key. And actually, you can set up a
1:34:28
hook in Cloud Code such that when it reads a secret, it tells you to rotate the key. So it's
1:34:34
kind of the responsibility of every developer to make sure the keys are not exposed and they
1:34:40
rotate them. But generally, like, for production secrets, don't just keep them in the file, like,
1:34:47
keep them someplace safe. So related to your question, Uzmain, like, with these tools,
1:34:54
we haven't had a huge incident happen yet. But it will be very important for everybody to increase,
1:34:59
like, their security processes and everything. And that just means, like, don't share passwords
1:35:04
via Slack, and try to have each employee have their own version of a secret, their own login,
1:35:10
don't share logins or anything like that. So that's part of the solution. All right, as you can see,
1:35:15
it's, you know, it's getting ready to use WhatsApp to send myself the message.
1:35:24
Okay, it's started typing. Sorry.
1:35:29
You can see it's typing, hopefully.
1:35:44
Oh, sorry, if I didn't get to your question, please feel free to just re-ask your question.
1:35:47
There's a lot, there's a lot of them here, but just feel free to repost it so that it's fresh.
1:35:53
Repost it so that it's fresh. Okay, so it did it. And you can see that this is something that
1:36:01
I can see on my phone before I start looking on my computer.
1:36:06
So I'm going to start the next task, but probably won't get to finish it,
1:36:10
but at least we'll have surveyed everything. So here's what we want to do next. We want
1:36:19
to build a routine that looks at our Google calendar, looks at the non-declined meetings
1:36:31
for external people, people outside the organization, looks them up on LinkedIn.
1:36:37
Then it looks at our Gmail and Slack for recent threads on the meeting. And then it sends ourself
1:36:45
an email, our daily brief email, so we can focus on our day, our meeting for the day,
1:36:51
and the most important information. So we've kind of already done all this, so we can start with the
1:37:03
prompt. So I'm going to put today's date in here. This is not crucial, but just to make sure the
1:37:13
agent is kind of aligned on the target dates. And actually, this is tomorrow, so this is going to be
1:37:20
May 1st. And what I've just described, that's what this prompt does. Yeah. And this is just a,
1:37:32
hey, if Slack's not working, just say, hey, Slack's not working. All right, so this is a
1:37:37
long detail prompt, but essentially it just does what I showed you in the diagram,
1:37:41
tries to create a daily brief email. Okay. So I'm going to do this in a new session, and
1:37:49
I'll choose the project that we had enabled before, the demo project.
1:37:54
So I've pasted it in, and I want to see it start working before,
1:38:00
so I have confidence that it's on the right track.
1:38:02
And yeah, we're going to see it start to do just this.
1:38:19
Okay. So it's doing this.
1:38:26
We just want to see that it opens Chrome. That will tell us everything we need to know,
1:38:30
that it's on the right track. So something to note here is that
1:38:38
Cloud creates its own to-do list. Before it starts on complex pieces of work,
1:38:44
it creates its own to-do list. And this is very helpful to look at, so you can see, hey,
1:38:50
what is it going to do? Is it going to do what I asked it to, or is it going to be off in the same
1:38:54
way? Okay. So now it's started working. So it's looking at my calendar of what meeting,
1:39:02
so it's going to pull that up. So one thing to note here is what I just described,
1:39:10
a human can do it much faster than Cloud. As you see here, Cloud is stopping. It's
1:39:18
stopping almost to understand the page before it decides what to do next.
1:39:23
It's looking at my appointment schedule and things like that here. A human can do this a lot
1:39:28
faster. So this kind of thing is still the place where humans can probably do this faster if you
1:39:36
just want to do it one time. But if you want to do it regularly, this is where a workflow like this
1:39:42
has value because you wouldn't have to do this yourself every time. Cloud would just do it and
1:39:48
you would just wake up to your daily brief email. Okay. Let me come back to the other agent that
1:39:59
we're working on.
1:40:00
And yeah, you know, essentially, I was able to send myself an email.
1:40:07
So we can just be like, great, update the schedule task with this WhatsApp instruction.
1:40:17
So just in natural language, no clicking any special buttons, you can just update your
1:40:22
schedule to do this additional action.
1:40:25
And I want to say, this is the best way to build a skill or a routine.
1:40:33
You want to demonstrate it to the agent first, then ask it to automate it, then ask it to
1:40:37
turn it into a skill or ask it to turn it into a schedule.
1:40:41
So yeah, you know, if we were to let this keep going, like it's just running through
1:40:46
this again, so that it can be a schedule application that can just run every morning.
1:40:52
Okay, so I'm not going to let it through because we're working on something else.
1:40:56
Okay, great.
1:40:57
So something I wanted to make you aware about with these routines feature is if you want
1:41:05
to do work overnight, this routines feature is how to do it.
1:41:11
You know, you essentially just describe like every hour, look, you know, look at a to-do
1:41:19
list that you created, and pick the next thing to work on and build it.
1:41:30
So as you can imagine, okay, this is like, this is a representation of, of like work
1:41:35
you can describe.
1:41:36
I want you to, as part of the homework for this is, think of what you can get to run
1:41:44
over the weekend, not just overnight, but like, just looking at what you have to do
1:41:49
next week, think of a research task that would benefit from a long session of web research
1:41:56
and maybe you have documents and things.
1:41:59
And this is what the prompt would look like to sell off the agents on this task.
1:42:06
So you know, I'm not going to create it now, but like, the most granular that Cloud lets
1:42:12
you get is an every hour.
1:42:15
You know, if you're, if you're a developer and you're using Cloud Code, this is, you
1:42:21
know, you can do finer than an hour, but as of right now, you know, Entropiq just lets
1:42:28
you do at most an hour.
1:42:30
And usually this is enough for like research activities that are happening overnight.
1:42:36
So I have example, you know, candidate workflows.
1:42:43
So we worked on a daily debrief.
1:42:44
You can work on an evening debrief.
1:42:47
If you're, you know, you're working at a NGO and you want to kind of like scan different
1:42:51
fundraising opportunities and the deadlines and, and, and see like, just build a brief
1:42:57
related site, you can do that, or you can just come up with something else.
1:43:00
Sorry, I'm rushing through this because we're running out of time, but as you get used to
1:43:06
doing workflows in Cloud Code and, and, and, and you want to graduate to agents, this is
1:43:13
practically how you will create an agent.
1:43:17
And in the prompt is how you will give it various levels of agency, right?
1:43:22
If you just wanted to do something specific and you've articulated it in a doc, you just
1:43:25
need to tell it like, Hey, let's look at this to-do list, long research to-do list, research
1:43:30
everything one at a time, you know, but you can just like say, Hey, we, we completed everything.
1:43:38
This project is almost ready, find a bug and fix it every hour or, or find a small
1:43:45
improvement and build it every hour that this is where you've given the agent more agency
1:43:52
now, you know, and that's also where it gets more dangerous, you know, because the more
1:43:57
agency you give it, the more stuff it can do in your computer.
1:43:59
So, you know, save this, this is fine for experimentation and research tasks, but truly
1:44:06
save, you know, the workflows that change things and that might break things for when
1:44:12
you fully feel confident.
1:44:13
All right.
1:44:14
So I'm going to take a couple of questions.
1:44:16
Are you able to do this with emails?
1:44:18
Yes.
1:44:19
Yes.
1:44:20
You can, as long as you can describe the automation and natural language, you can fully automate
1:44:25
emails.
1:44:26
Um, you know, I, I, I'm using Chrome, but if you set up the Gmail connector, these workflows
1:44:35
can be even faster so that it can happen throughout the day.
1:44:38
Like, Oh, you know, when you get an email, you know, with a question that's in your,
1:44:42
your internally maintained frequently asked questions list, the AI can just auto respond,
1:44:48
you know, that that's something that can happen.
1:44:49
And you would set it up by using the schedule tasks.
1:44:52
Like, again, the most granular we can get is every hour, but, you know, if you want
1:44:58
something more fine grained, that's when you will look into cloud code, you know, you feel
1:45:03
like ready for more leverage, more automation, you can look into cloud code.
1:45:08
So there's that.
1:45:09
Um, okay.
1:45:10
So, all right.
1:45:11
So, you know, we started working on this for some reason, but the, the, the PowerPoint
1:45:18
file is ready now.
1:45:21
So I'm actually going to open this in a Google slides, just so I can take a better look at
1:45:26
it.
1:45:27
Okay.
1:45:28
Um,
1:45:43
Okay.
1:45:58
Okay, so it created the PowerPoint slide, and we got full fidelity when moving it to Google Slides. There's no like, um, things that are off, you know, and even if there's things that are off, where it's natively created the PowerPoint slide that is fully editable, you know, so there's some first finishing policies, we don't need to add for.
1:46:26
But again, you know, most of the media changes that you would want to make, you'd want to do it in chats, you know, like when you're chatting with the coworker versus like, waiting here, you know, you kind of want to do the final post stuff here, maybe like if things are a little off, or, you know, go back to cloud to get it to change. So yeah, you know, as you build workflows and build an opinion, you can build branded documents and things.
1:46:56
Okay, so so we're running short on time. So I want to skip ahead to the reflection part.
1:47:26
They kind of get impatience when when creating these prompts, so they rather just type in short. So we build this tool called a fieldmark that is such that as you're prompting, if you just have, you know, an idea for what you want to prompt, but the next output is so important, you want to like make sure you have the best prompt possible, you can actually use AI to help you generate the next best prompt.
1:47:51
So I'll say like, for this, you know, WhatsApp, WhatsApp myself a schedule for the day for how to dress and stuff like that. I'll just say like, how do I make this faster, you know, to to time to write a super detailed prompt.
1:48:11
So we're gonna copy over this prompt coach prompt that tells the agent like, hey, look at this agency fun site with all the guidance about how to build a better prompt baked in. And it's going to read that and it has the understanding of the context to generate the next best prompt. So we should see that soon.
1:48:34
And, okay.
1:48:36
And in the meantime, you know, if you can offer feedback, we would try to cram a lot here, where this ideally should be a longer workshop and a bit more bespoke depending on the org, we try to show you many different things from creating a design system that you can share with your colleagues, so that they can make slides in the same fashion, consistent fashion across the org, to building the next best prompt.
1:49:03
To building a daily brief agents and scheduling workflows so that work can happen while you're away from your computer.
1:49:11
But, you know, there's a much more we could have explored.
1:49:15
If you can fill out this feedback, it would be much appreciated.
1:49:19
And we're having a real meetup, you know, to go into the codecs app next.
1:49:26
So, I can preview the codecs app here, if you can see my screen.
1:49:29
Just to give you a little preview.
1:49:32
Where Claude has this opinion about non-technical users and technical users, non-technical users use code, codecs doesn't have that distinction.
1:49:45
But it is pretty easy to use as well.
1:49:50
If you have something you want to do, like, hey, show me the sites, you know, something like that.
1:49:58
This is for the agents and workflows site that I'm working on.
1:50:01
It has a built-in browser.
1:50:06
Because if you remember, OpenAI built a browser called Atlas.
1:50:11
And they've now tried to consolidate all the things that they've built into one platform.
1:50:18
So, yeah, you know, here it is.
1:50:21
Just something I want to demo real quick.
1:50:24
If you're not technical and you just want to be able to point out certain things, OpenAI makes it much easier to, like, add feedback on the UI.
1:50:33
Change this, you know.
1:50:35
You can just say, like, change this.
1:50:37
And you can imagine, like, as you're browsing the site, you have opinions about how to change and it can change it.
1:50:43
Yeah.
1:50:45
Okay, great.
1:50:47
I realize that we're out of time.
1:50:50
And we've demoed a lot.
1:50:52
So please leave feedback when you have a second.
1:50:56
Again, you can always visit this site.
1:50:58
I'll paste it in the follow-up to go through the course at your own pace.
1:51:03
Please feel free to reach out.
1:51:05
If you have any questions, I'm more than happy to, you know, give advice about these agentic workflow stuff.
1:51:12
You know, as a field, we're all learning about that stuff.
1:51:14
But it is the case that in the private sector, in the development sector, there are a few that are getting a lot of leverage, you know.
1:51:21
As one person, they're getting the work done from whole teams.
1:51:25
And, yeah, you know, I think it's important for all of us to explore that.
1:51:29
Thanks, everybody.
1:51:31
I really appreciate it.
1:51:32
And I'll see you next time.
1:51:34
Cool.
1:51:41
I'm still around.
1:51:42
So, you know, if I didn't get to answer your question, I'm more than happy to chat now.
1:51:54
Okay.
1:51:55
Great.
1:51:59
Yes.
1:52:07
Yes, yes.
1:52:08
The recording will be available after this call.
1:52:13
You should be able to unmute yourself and let me ask you a question.
1:52:33
Yes.
1:52:35
Thank you so much.
1:52:37
I'm having some trouble, like, allowing my cloud to access WhatsApp.
1:52:45
It says there's no connector available.
1:52:48
I know it's just a small bit of what you did, but I felt that it would be quite useful to sometimes be able to send it back to myself.
1:52:57
Do you know why it might be saying that?
1:53:01
You have the cloud desktop app installed?
1:53:04
Yes.
1:53:06
Okay.
1:53:07
So explicitly say, like, computer use.
1:53:10
So it's saying that, oh, you haven't given me access to a WhatsApp connector.
1:53:16
And there are those that exist, but you wanted to use computer use.
1:53:20
You essentially wanted to take over your computer and do it.
1:53:24
You may have to enable certain permission, if you can see my screen.
1:53:31
Yeah, yeah.
1:53:33
Yeah, yeah.
1:53:34
There's a computer use connector.
1:53:36
So you may have to go in cloud settings.
1:53:43
Let's see.
1:53:45
Okay, so this answers, I think.
1:53:47
I can look in there.
1:53:48
Thank you so much.
1:53:49
I appreciate it and appreciate this session today.
1:53:53
Yeah, no problem.
1:54:02
Yes.
1:54:04
Sorry.
1:54:08
Raspus, is that you?
1:54:11
That is me, yes.
1:54:13
Thank you so much.
1:54:14
Great.
1:54:15
And I just sent a note also.
1:54:17
So I'm the board chair of a platform.
1:54:21
We're an NGO as well, but called Epic Africa.
1:54:23
We have more than 5,000 African CSOs on it.
1:54:27
We have like a LinkedIn for African CSOs and then the back end also for funders
1:54:30
so they can identify CSOs on the ground instead of all these, you know,
1:54:35
Western-based intermediaries.
1:54:37
And we've been really wanting to do something on AI for African NGOs on a webinar.
1:54:43
We do a bunch of these.
1:54:44
We usually get 500, 600 people on them.
1:54:47
So just in case you're interested, we'd love to chat about that.
1:54:51
And I'm happy to put you in touch with Rose, who is our CEO, who's Kenyan,
1:54:57
but based in Dakar.
1:55:00
So, yeah, should I send you over an email?
1:55:02
Yes.
1:55:04
Yes, definitely.
1:55:05
Yes, we'd love to connect and figure out, yeah, how we can plug into that.
1:55:10
We want to reach as many people as possible, especially since it's so early,
1:55:13
you know.
1:55:14
Yeah.
1:55:15
And these questions are so valuable, you know,
1:55:18
because it helps me understand, like, what people are interested in working on.
1:55:22
Yeah, yeah, yeah.
1:55:23
And so, you know, the great thing is that we're sort of a, like a,
1:55:26
we have a lot of distribution and our value to those CSOs is really, like,
1:55:31
getting them value.
1:55:32
And so this is a question we get all the time.
1:55:34
We've done some AI things.
1:55:36
So I'll shoot you over an email and I'll connect.
1:55:38
I'll add Rose in as well.
1:55:39
Thank you so much again.
1:55:40
This was really great.
1:55:41
And have a great evening.
1:55:42
Okay.
1:55:43
Thanks, Ross.
1:55:44
I appreciate it.
1:55:45
Thanks, man.
1:55:47
Andrew.
1:55:48
Yeah.
1:55:49
Hey.
1:55:50
Great.
1:55:51
Thanks, Edmund.
1:55:52
This is really helpful to have just bespoke questions and answers and all
1:55:56
that.
1:55:57
And it's lovely to see how you work and all the stuff that agency fund has
1:56:01
built out.
1:56:02
So, you know,
1:56:03
I'm struggling a little bit because I've done much of the same in chat GPT,
1:56:07
and it's not quite as orchestrated or well orchestrated as I've seen people
1:56:12
do this in Cloud.
1:56:14
And I'm eager to move over, but part of me,
1:56:16
particularly the issue around bumping up against, you know, tokens,
1:56:21
credit limits or costs, as well as, you know,
1:56:25
there are some limitations and you pointed out like generating PNGs and
1:56:29
things like that.
1:56:30
There's always a use for using other models.
1:56:32
So two questions.
1:56:33
One, I'm worried that, you know,
1:56:36
the advancement in one model is then superseded by another.
1:56:39
You're constantly bouncing back and forth.
1:56:41
And that's super frustrating because you kind of have to choose one,
1:56:44
at least at some point, right.
1:56:46
Or at least for some time.
1:56:47
And I feel like the switching cost of switch is high.
1:56:50
And then the second is that, you know,
1:56:53
is there not an orchestrating environment,
1:56:55
like an open work or orchestrating environment where you can actually build
1:56:59
out consistent across like, you know,
1:57:02
you can swap out the models as needed rather than swap the orchestration
1:57:08
space as needed.
1:57:11
Yeah. Yeah.
1:57:13
So what you're describing the ability to basically just invest in one,
1:57:18
you know, app, learn, learn that app,
1:57:21
but able to pull in different models for the right task.
1:57:24
Yeah. And right now the best option for that is cursor.
1:57:28
So cursor is this, it's basically a developer IDE.
1:57:34
Yeah. And I would say.
1:57:36
I've heard of it. Yeah.
1:57:39
Yeah. You know, if you look at the interface,
1:57:41
it looks very similar to what we're doing.
1:57:43
Like, you know, you're opening files.
1:57:45
It has a built-in browser, but it's heavily oriented for developers to use.
1:57:51
It's very heavily oriented for developers to use.
1:57:54
But on this situation, this, you know,
1:57:59
this Game of Thrones style situation where one model is on top the next
1:58:02
week, the next, this other model is on top.
1:58:05
Yeah.
1:58:06
There's not a great solution for that other than, you know,
1:58:10
betting on one of the two top models, which is a cloud or a codex.
1:58:15
So yeah.
1:58:17
When each of these release a new feature, my cloud design,
1:58:21
I would say a month there's a probationary period where this,
1:58:25
whether this feature might be even alive.
1:58:27
So I wouldn't rush into like considering switching that because this org
1:58:31
released that and that org released that because they will catch up,
1:58:35
you know, they will catch up to each other where there are gaps, you know,
1:58:38
I just have to give it like a few weeks and,
1:58:40
and there's a lot of churn right now in the feature set.
1:58:43
But, you know, if you have workflows that, you know,
1:58:47
if you just want to be able to, you know,
1:58:51
especially on this more intensive stuff, like, you know,
1:58:54
doing stuff overnight and then things like that, I think is better to,
1:58:59
right now cloud is in the lead, you know, cloud is in the lead.
1:59:03
Where I fall short is this usage stuff that when you start using cloud and
1:59:09
like, you're just using the best model for everything,
1:59:11
you're going to run out of tokens, you know?
1:59:13
And then if you started relying on it, like, you know, you're gonna, yeah.
1:59:17
You know, it's going to feel awkward to like,
1:59:19
well how do I work like without this thing?
1:59:22
So that's why I suggest using Sonnet by default.
1:59:27
If you learn this one tactic of using Sonnet by default and only pulling in
1:59:32
Opus where you feel like this is important, this isn't,
1:59:34
I want the most intelligent model, you're never going to run out of tokens.
1:59:38
You're never going to run out of tokens.
1:59:39
You're never going to run into this issue that many other people who are just
1:59:43
using the best model for every little thing to like, you know,
1:59:46
change a sentence and stuff like that.
1:59:49
So while Codex right now, Codex right now is like, you know,
1:59:53
giving out these, yeah, they're giving out these promotions,
1:59:57
like resetting the usage.
2:00:00
It feels like you have so much usage,
2:00:02
but at the end of the day, at both orgs,
2:00:04
these tokens are being subsidized.
2:00:06
So there will come a point where we all need
2:00:08
to be more cost-effective in how we use these tools.
2:00:12
So I would say it's safe to bet on Cloud for now.
2:00:15
If you're on OpenAI and you can wait,
2:00:17
I think generally, yeah, I think you can wait.
2:00:21
Just give it a month and all the features that Cloud has,
2:00:25
OpenCodex will catch up.
2:00:26
Specifically, the Codex desktop app.
2:00:28
If you're doing things in charge of bt.com,
2:00:31
you're really limited in what you can do
2:00:33
because it doesn't have access to your computer.
2:00:35
But if there's a platform you want to start to invest in,
2:00:39
but still stay within OpenAI, it's Codex.
2:00:42
The Codex desktop app, that's this thing.
2:00:44
That's what it looks like.
2:00:47
This is the big thing that I showed you.
2:00:55
Yeah, how do you manage token limits
2:00:56
when Cloud is working on automating
2:00:59
in the backend for daily repeat tests?
2:01:00
Yeah, yeah, this is a great question.
2:01:03
So you're not going to run into token limits
2:01:07
because of the hourly cadence that Cloud says
2:01:13
that you can't do any tasks
2:01:16
with a bigger granularity than an hour.
2:01:20
You can't do something every minute.
2:01:22
They intentionally put this because they just don't want,
2:01:24
they don't have the servers to serve all these computes.
2:01:27
So just in using this feature, you're not going to run out
2:01:31
unless you're just creating so many scheduled tasks.
2:01:34
You can imagine you start to use this
2:01:36
and you have 10 things happening every day.
2:01:41
With using these tools, a good principle to have
2:01:44
is make sure you call things that you're not using.
2:01:50
Yeah, you can get into the mode
2:01:52
where you're having so much happening,
2:01:54
but you should be always asking
2:01:56
what's actually providing you a value?
2:02:01
What's providing you a value
2:02:02
and then just cut or combine things that are not.
2:02:05
But again, like I said, if you make this run
2:02:10
in a certain model, like Sonnet or something like that,
2:02:14
you can configure what model you want it to run.
2:02:18
If you make it run on a non-Opus model, basically,
2:02:21
like Sonnet, you're not going to run out of tokens.
2:02:24
You can get to the point I've seen like running,
2:02:26
like I think 50 routines was the limit.
2:02:31
I'm not sure if there's still a limit now,
2:02:32
but when I was using it,
2:02:35
the maximum number of scheduled tasks
2:02:37
you could have at any one time was 50
2:02:39
and you're not going to run out of tokens at all.
2:02:41
It's only when you're using Opus for every single thing
2:02:44
that you need to worry about tokens.
2:02:47
And like I talked about didactically,
2:02:49
like for educational purposes, when you get started,
2:02:51
yeah, use Opus just to get value so you can see these things
2:02:55
but you should soon think about a default of Sonnet
2:02:59
and then pulling in Opus when you need it.
2:03:03
Is there a question, Shruti or Rasmus?
2:03:12
Can you hear me?
2:03:14
Yes, I can hear you.
2:03:15
Okay.
2:03:16
Hey, so I'm the founder of an AI studio,
2:03:20
an ecosystem builder, it's called Tilted Ground
2:03:23
and I work with nonprofits in India,
2:03:27
nonprofits and funders.
2:03:29
And I'm just, I'm actually, I love the workshop today
2:03:33
and I would love to keep attending more of the workshops.
2:03:37
I have a question more around like,
2:03:40
how do you see working with,
2:03:44
kind of in the format of these workshops?
2:03:46
Do you intend to support more technical users
2:03:51
or more non-technical users?
2:03:53
And do you intend to focus on specific AI tools
2:03:57
in these workshops?
2:03:58
Or do you also see yourselves consulting
2:04:01
and working with organizations
2:04:04
more in a consultation capacity as well?
2:04:09
Yes, good question.
2:04:10
So one on, yeah, we already in a consultative capacity
2:04:17
with technical engineers in our portfolio,
2:04:19
we already do that where we're like,
2:04:22
we're jumping into a code base,
2:04:23
we're showing them, hey, let's use cloud code
2:04:26
for code review, let's set up automation
2:04:31
to make it easier, make it easy to build things quickly.
2:04:35
So that's something we already do with our portfolio.
2:04:38
With this content, I think we wanna like start
2:04:43
with the non-technical audience
2:04:44
because we feel like maybe they look at AI
2:04:49
as like, they're not sure like what the kind of things
2:04:52
that they can do with AI.
2:04:54
And at our organization, at the range of organizations,
2:04:59
they often don't have a single developer,
2:05:01
they don't have a single developer.
2:05:02
So it's gonna be a non-technical person
2:05:04
that is setting this all up and showing their coworkers
2:05:09
how to do it.
2:05:10
So with the web content, the self-paced content,
2:05:13
I think that's where we're gonna end up.
2:05:14
But we're also definitely interested
2:05:15
in like jumping into an org.
2:05:18
This is something that we're increasingly have
2:05:21
more and more conversations about,
2:05:23
like jumping into an org, like the way we kind of do now,
2:05:28
but now we focus on like making them
2:05:32
a learning organization, making them run A-B tests
2:05:34
and getting the data engineering up.
2:05:36
But also like getting them to be-
2:05:39
You're gonna need to file an email about the agency file.
2:05:43
Getting them to be AI native,
2:05:44
like understanding what is the work everybody does
2:05:48
and helping them convert that to more AI native,
2:05:51
faster workflows over time.
2:05:53
So that's definitely something we're interested in.
2:05:56
And yeah, feel free to reach out if you have an opportunity.
2:06:02
Sure, that'd be great.
2:06:04
And actually one specific area
2:06:05
that I'm looking to engage deeper on
2:06:08
is data readiness in organizations.
2:06:12
Because, you know, especially like even looking
2:06:14
at doing any kind of experimentation work with nonprofits,
2:06:18
it really requires data to be in a certain format.
2:06:22
So if that's something you may be interested
2:06:23
to talk more about, I'd love to reach out
2:06:25
and get your thoughts.
2:06:28
Yes, definitely.
2:06:30
Yeah, just admin that agency file.
2:06:32
I think I pasted it.
2:06:33
But yeah, feel free to reach out anytime.
2:06:36
Awesome, thank you.
2:06:41
Cool.
2:06:43
Yeah.
2:06:44
Any other questions, please feel free to raise your hand
2:06:48
so I can unmute you.
2:06:53
Meantime, let me look at this LinkedIn post.
2:07:08
Yeah, yeah.
2:07:09
So he's,
2:07:13
yeah, so this is a good point.
2:07:15
Like, you know, people are already like, wow, you know,
2:07:17
if Entropic is, and then OpenAI is like sponsoring
2:07:22
thousands, tens of thousands of usage,
2:07:25
can I even rely on this tool?
2:07:28
Like if you had to, if one person using this tool
2:07:32
like meant an $8,000 a month bill in your org,
2:07:37
is it still worth it?
2:07:39
You know, that's a real question.
2:07:43
But here's what I'll say is promising.
2:07:47
It's getting easier and easier to run the open models
2:07:50
on your computer.
2:07:51
I mentioned like, you know, right now,
2:07:52
if you want to do it confidently,
2:07:55
you need like, you need like 64 gigs of RAM computer,
2:07:59
you know, average computer is like eight gigs of RAM.
2:08:03
But it's getting easier and easier to run
2:08:05
on your own computer.
2:08:07
And you see the way I propose using Sonnet for most things,
2:08:12
and then only pulling in the frontier models
2:08:14
when you need it.
2:08:15
I think that's the future we're headed towards.
2:08:18
And in fact, I saw that open cloud is gonna support
2:08:22
the ability to use open models within the cloud desktop app.
2:08:29
So, you know, for example, you see where we're picking
2:08:34
between the cloud models here,
2:08:35
but they will soon add an option such that you could pick
2:08:38
like Lama or the Gemma models and things.
2:08:42
And that will radically change the cost structure
2:08:45
because most things we're doing with these AI,
2:08:48
you know, can be done by a cheaper, less intelligent model.
2:08:54
We only need to pull in the biggest models
2:08:58
when we're early, like in the planning phase.
2:09:00
And what we decide in the planning phase is crucial
2:09:04
for the rest of the work, you know, the implementation phase.
2:09:07
So, yeah, I feel like open models will get better.
2:09:10
It might mean like, you know,
2:09:12
making sure we have computers with more memory
2:09:15
or attaching memory or running things on the server,
2:09:18
but we won't just have to rely on open AI in cloud
2:09:23
if we wanna use these tools in our workflows in the future.
2:09:25
There will be a good enough open model,
2:09:29
open options, open source options.
2:09:32
Oh, yeah, this is a question, Karen.
2:09:38
Yeah.
2:09:39
Well, yeah, with that, if there are no more questions,
2:09:43
please feel free to reach out again.
2:09:47
There's no chat here, so it's hard to add my email,
2:09:49
but my email is admin at agency.fund.
2:09:53
We wanna do more of these, so if you can reach out,
2:09:55
just opportunities for collaboration
2:09:57
or, you know, content that we should add,
2:10:00
things we should focus on, it's for free.
2:10:02
Yes, Ana.
2:10:03
Ana, I just wanted to thank you.
2:10:05
Okay, oh, thanks.
2:10:07
Yes, with that, thanks, everybody.
2:10:10
Please, I hope you have a good rest of your week.
2:10:12
Bye.
2:10:32
Bye.
Keep exploring
May 2026
Learn Enough Codex To Be Dangerous
Mar 2026
AI 101
Jul 2026
Evaluating AI for Social Impact
Jun 2026
Behavioral Science at TAF
May 2026
AI Agent Evaluation Tutorial
Apr 2026
AI Evaluation For The Social Sector
Jun 2025
AI4GD 2025 Model Eval Workshop
Jun 2026
Early Deployment of an Integrated Digital Platform (shamiriOS) for Scalable Youth Mental Health Service Delivery in Kenya: Development and Usability Study
Mar 2026
Model Eval for Leaders
May 2026
GenAI for Data Analytics