Grok Bot has been very popular lately, and rightly so. I believe this is what an AI assistant product should look like, rather than a session-based chat with Codex or Claude Code. We’ve become too used to a chatbot experience built more around an agent’s context window than the way humans work. Today, let’s break down how Grok Bot is put together.
What’s in each bot’s context window
As you would imagine, each bot has its own context window. Grok Bot’s developers want us to think of each bot as a human teammate, and this is exactly what they had in mind when designing its context window.
The rule is simple: the earliest message a bot receives appears at the top of its context window, and the latest message appears at the end. It doesn’t matter how many channels the bot is in or how many other bots it is talking to. It always follows that rule. This is also better for prompt caching.
One context, across channels
1 / 5Find the papers we should read.
1 of 5 received. m01 from User in #research is at the end of the context.
Key Tools
There are other tools I won’t cover here. I’m only mentioning a few that I find interesting.
sendToUser
A bot can work, send several messages, or finish without sending anything. You can probably guess how: a sendToUser() tool. Normally, the final text response is the focus, but an agent can generate only one final response per turn. With this tool, a bot can send multiple messages to the user and provide updates as it goes.
Tool arguments:
| Type | Required payload | Optional fields |
|---|---|---|
text |
content |
images, reply_to, channel, to |
attachment |
url |
alt |
widget |
widget containing prompt and options |
Within widget: helpText, multiSelect, allowCustom, dismissOnMoveOn |
cursor-agent |
bcId |
None recorded |
secret-request |
secret containing label, connector, and field |
None recorded |
There’s no dedicated end-turn tool. Because the user can’t really see the final response, each bot’s system prompt tells it to keep that response short to save tokens. One-word final responses like “done” or “Answered” are common.
An alternative design would be to add a dedicated end_turn tool, but that would likely consume more context than a simple one-word final response.
sendToAgent
Agents can also talk to each other, just as they use sendToUser to talk to the user.
Tool arguments:
| Argument | Required? | Details |
|---|---|---|
target_id |
Yes | The target’s ID, not its name |
message |
Yes | The message to send |
images |
No | 1:1 only |
priority |
No | 1:1 only; wakes the recipient immediately |
ListAgents
Grok Bot doesn’t actually have this tool. All bots use a shared computer. Every time a bot is created, a directory is also created as its workspace, along with an entry in profile.json. Even without a ListAgents tool, a bot can discover other bots by inspecting profile.json on the shared computer.
Every time a bot receives a message, it “wakes up”: the latest message is appended to its context, and a new turn begins. The message also includes metadata such as its message ID, the sender’s agent ID or user identity, and the channel ID where the message appears. A green “online” indicator shows that the agent is currently working.
What comes next
The bot is the unit of continuity, but its context window is still finite. As messages accumulate across channels, I think compaction will become a big part of making this design work over time. What gets retained, summarized, or dropped will shape how useful the bot remains. Compaction alone won’t be enough: automatically saving useful logs and reusable skills in the bot’s workspace is essential, so it can recover past decisions and reuse what it has learned without keeping everything in context. To me, that’s the next challenge: not just making a bot feel like a teammate in the moment, but giving it continuity after its context has been compacted.