Intro
You start a long AI research task, close your laptop, and the run dies with the app. That is the failure mode AWS Pizza Bot self-hosted is built to fix. According to the MarkTechPost announcement (opens in new tab), the tool started inside Amazon, served more than 2,000 people for meeting preparation, email drafting, Slack summaries, CRM logging, and research, and then got rebuilt as an Apache 2.0 open source project with macOS, Windows, and Linux desktop builds plus browser and terminal clients.
The distinction that matters in this release is the embedded desktop server versus an always-on backend. Quitting the desktop app stops its embedded server and ends active runs. LangGraph checkpoints preserve the thread state, but the step in flight can be lost. An always-on backend keeps scheduling and execution running after the client disconnects, and reconnecting clients can replay buffered events. That is the difference between a tool that feels like a toy and one that behaves like infrastructure.
I run a small VPS hosting business, so I have spent years thinking about what keeps running after a customer closes their browser or laptop. A cron job on a shared host, a queue worker processing jobs, a database connection pool. The principle is the same here as it is for Pizza Bot: if the process is tied to a client session, the session ending kills the work. If it is on a server that stays up, the work finishes and the result waits in the inbox for you.
That is the lens I am bringing to this post. The desktop builds are convenient, but the self-hosted backend is the part that matters if you want AI tasks to survive you quitting the client. The rest of this post covers where Pizza Bot came from, what the self-hosted deployment actually requires, and how to test the difference yourself in a few minutes.
Background: Where Pizza Bot Came From
Pizza Bot started as an internal Amazon tool. It handled meeting preparation, email drafting, Slack summaries, CRM logging, and research for more than 2,000 people inside the company. The public release, covered in the Marktechpost announcement (opens in new tab), was rebuilt as an Apache 2.0 open source project. That rebuild matters. This is not a weekend side project that AWS threw over the wall with a blog post and a promise.
The runtime uses DeepAgents and LangGraph for stateful execution. A Hono API server owns runtime execution and storage, and SQLite holds cross-thread memory and application metadata. Clients talk to the server over HTTP and server-sent events. The architecture is the unglamorous part of the announcement, but it is the part that decides whether your long-running task survives a disconnect. LangGraph checkpoints retain thread state and approval pauses. Reconnecting clients can replay buffered events. Closing a thread or disconnecting a client does not stop a running server.
I have seen this pattern before in my own hosting work. A customer runs a long database migration from their laptop, the Wi-Fi drops at a coffee shop, and the migration dies halfway through. The fix is always the same: move the work to a process that is not tied to a human being’s session. Pizza Bot’s internal history suggests AWS learned that lesson years ago. The tool was built for people who close their laptops at the end of a meeting and expect the Slack summary to be waiting when they get back.
The interesting part for self-hosters is that the public rebuild keeps that server model. You can run the backend on a small VPS and connect from any client. My expectation, and this is my own read rather than anything AWS promised, is that the standalone backend will become the default way people run this in production, with the desktop embedded server remaining a demo path. Keep that in mind when you read the deployment details next: the server is the product, the desktop app is just a window into it.
What’s happening now: AWS Pizza Bot self-hosted vs the desktop server
The core distinction in the source article (opens in new tab) is between the embedded desktop server and a standalone backend. The desktop builds ship with a server baked into the app. That is convenient for trying the tool, but it has the same weakness as running a database migration from your laptop. Quit the app and the server dies with it. LangGraph checkpoints preserve the thread state, so conversation history and approval pauses survive, but the step in flight can be lost. If your agent was mid-call to a model or mid-write to a file, that work is gone.
The standalone backend fixes that. You run it on a server that stays up, point your desktop, browser, or terminal client at it over HTTP, and the client becomes a window into work that lives elsewhere. Closing the desktop app does not stop the backend. Reconnecting clients can replay buffered events, so you do not miss the updates that happened while you were away.
Scheduling is where the difference shows up in everyday use. The server owns cron triggers, and trigger occurrences are recorded durably. If the server is down when a scheduled run was supposed to fire, it produces one catch-up run instead of replaying every missed interval. Think of it like a backup job that missed its 2 a.m. window. You want it to run once when the server comes back, not fire ten times to make up for ten missed nights. The design here matches that expectation.
For a self-hoster, the practical question is whether you trust a process that lives inside a GUI app. I have had to explain to hosting customers why their background job died when they closed their browser tab. The answer was always the same: the job was a child of that session. With AWS Pizza Bot self-hosted, the standalone backend separates the two. That is the version worth deploying if you want tasks to survive you quitting the client.
What it means in practice: AWS Pizza Bot self-hosted deployment
The standalone backend is not heavy. A single SQLite-backed Pizza Bot backend runs fine on a small VPS, the kind with 1 or 2 GB of RAM that I sell to hosting customers for about six dollars a month. The constraint is that each SQLite data directory supports only one backend process, so if you want multiple isolated backends you need multiple data directories and probably multiple servers. For a solo operator running a handful of skills, one small server is enough.
Model providers are configured under Settings > Providers. The list covers Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter, and local models through Ollama. For a self-hoster, Ollama is the interesting one. It lets you run the whole stack on hardware you control and keep task data away from external APIs entirely. If you are feeding the agent internal CRM notes or Slack summaries, that distinction matters more than any feature in the UI.
Approval controls are real but not automatic. A skill author must configure interruptOn and allowedDecisions on the relevant tools before you see the approve, edit, or reject buttons. If a skill does not declare those policies, the agent runs without pausing for approval. Before you connect a skill to a tool that writes to production data, open the SKILL.md and check whether the author actually gated the action. I have seen this pattern go wrong enough times in other tools to say it plainly: an approval UI is worthless if nothing triggers it.
The sandbox is where the design earns trust. The agent’s JavaScript interpreter has no network access and no host-filesystem access, and the filesystem layer works through explicit folder grants. That separation is the difference between an agent that reads only what you handed it and one that can wander your home directory. The project is Apache 2.0, as the source notes (opens in new tab), so you can read the sandbox code yourself before connecting your own data. That is more than you get with most hosted agent products.
What to expect next
The SKILL.md pattern is the part I expect to drive most of the ecosystem growth. Each skill bundles instructions plus scoped tool access, which means the community can package a worker the same way we package a Docker image or a cPanel plugin. I expect MCP integrations to multiply quickly, especially the Claude Code-compatible .mcp.json support, because that lowers the barrier for anyone who already has tools configured for another agent.
Ollama support will probably be the entry point for self-hosters. I say that because the people who want a background agent on their own infrastructure are usually the same people who do not want their CRM notes or Slack exports flowing through someone else’s API. Running the whole stack on a small VPS with Ollama is the cheapest way to test whether Pizza Bot fits your workflow. It is also the easiest way to get burned if you skip the approval configuration, but that is true of any agent tool.
The one-backend-per-SQLite-directory limit is the first pain point I expect people to hit. It is fine for a single server and a handful of users, which is probably where most self-hosters start. But the moment you want horizontal scaling or a second backend process for isolation, SQLite becomes the constraint you have to design around. I have hit the same wall with other single-writer databases, and the fix is usually moving to Postgres, which means the project will need a storage abstraction layer before that day comes.
The catch-up cron behavior is the thing I am least sure about. Recording trigger occurrences durably and collapsing missed intervals into one run is a sensible design, and I have seen similar patterns work in job queues. But this project is young, and I do not know how that behavior holds up when a server comes back after a long outage with dozens of schedules due. I would test that scenario on a throwaway VPS before I trusted it for customer-facing work.
My expectation, and this is a guess rather than a promise, is that Pizza Bot settles into the same niche as self-hosted tools like n8n or Windmill. The value is not the agent itself. The value is that you can point it at your own data and your own servers and leave it running after you close the laptop.
Your Next Step: Run the Interactive Explainer
The fastest way to understand what always-on does for you is to run the interactive explainer (opens in new tab) from the article that announced Pizza Bot, before you install anything. It is a simulation, so no external calls happen, but it models the exact failure mode that matters. Set the backend location to “Always-on server”, start the research brief task, then close the desktop client while the run is in progress. Reopen it and check the inbox. The result is there, waiting in Unread, because the backend kept executing after the client disconnected.
Then switch the backend location to “Embedded in desktop” and repeat the same run. Close the client mid-task, reopen it, and you will see the difference immediately. The thread is preserved by the LangGraph checkpoint, but the step that was in flight is gone. For a site owner this is the same lesson as closing your browser tab in the middle of a cPanel backup. The job was running inside that tab, and it died with it.
The simulation only takes a few minutes, and it shows you the approval flow too. The custom skill gates its publish tool with interruptOn, so the run stops at the proposed action and you get to approve, edit, or reject it. That is the moment to check whether the approval controls are actually useful for you. If you are going to connect this to real data, you want to know what a paused action looks like before it happens on a live system.
When you are ready to move past the simulation, the cheapest real setup is Ollama on a small VPS with a single SQLite directory. You do not need a GPU-sized instance for a test. A single-core box with a few gigabytes of RAM will run a small model slowly, which is fine for learning the workflow. Create a skill that has interruptOn configured for its main tool, start a task, and close your laptop. Leave it closed for an hour.
When you come back, open the inbox. If the result is sitting in Unread, you have a background agent. If the thread is stuck at the checkpoint with the step lost, you have the backend location wrong and you know exactly where to look. Either way, you have spent less time than it takes to read the GitHub README, and you will have a much clearer idea of whether Pizza Bot belongs on your server.