Intro: Why “we’ll just sanitize it” isn’t enough anymore
The deliverable has changed. It used to be a PDF or a static report. Now, for a growing number of platforms, it is an interactive HTML file generated by an AI model. The shift from human-authored to machine-authored code changes the security calculus fundamentally. You are no longer trusting a colleague; you are executing a program created by a probabilistic system trained on the entire internet. The core problem is running this untrusted code inside a browser tab that is authenticated to your platform, holds session cookies, and can access sensitive client data. The old approach of upstream sanitization, checking the file for known-bad patterns before it ever gets served, breaks down. Sanitizing scripts destroys the interactive feature that made the HTML valuable in the first place. So, a new guiding principle is needed: nothing upstream of the browser gets trusted. Not the uploader, not the prompt, not your own quality assurance step. The security boundary is where the code meets the viewer’s browser.
Background: Building a hosting platform for untrusted HTML
I recently dealt with a similar product requirement for a small VPS hosting client. They needed to let administrators publish self-contained, interactive HTML dashboards to their clients. The initial ask was simple: upload a file, get a revocable link. The first realization was that sanitizing JavaScript away would kill the product’s purpose. The more dangerous realization was the risk model. An uploaded HTML file is a program. It runs with the privileges of the page it lands on. If that page is your authenticated application, the script can read cookies, make API calls as the user, and exfiltrate data. The specific, and entirely plausible, attack chain is prompt injection. Someone uploads a document into a chat interface; that document contains text aimed at the AI model, not the human reader. When the model later generates an HTML artifact, it follows that injected instruction, embedding malicious JavaScript. That script then executes in the browser of a client-side stakeholder who clicks the link. The author might not even realize what they uploaded. Trusting the author is no longer a viable security model when the author is an AI agent.
AI-generated HTML security: The three browser boundaries that matter
Mitigations around the generation process are helpful engineering. Isolating the agent’s context, requiring self-contained output, and running an automated check on the result all reduce risk. But they are not security boundaries. A determined or sufficiently clever piece of injected content can bypass them. What actually contains the risk are the boundaries enforced by the viewer’s own browser. The real security model is not about what you do on your server before sending the file; it is about what the browser allows the file to do once it arrives. This splits into three independent pillars: identity, origin, and capability. Breaking one of these does not automatically compromise the others. A leaked URL still checks identity. A malicious script runs in a sandboxed origin. Code that gets through both still lacks the capability to reach sensitive APIs. This is the model described by Islam Gagiev in Three Security Boundaries for Hosting Untrusted, AI-Generated HTML (opens in new tab), and it is the model any platform serving such content must adopt.
What is the identity-origin-capability model?
Think of these as three separate locks on the same door, each using a different key. Identity is the first check, enforced by your server. It answers: is this person allowed to see this artifact? This is your standard platform authentication and access control. The link is not the credential. A short-lived token in the URL is a bearer credential that can be forwarded, which is bad. Instead, every request to view the artifact must pass through an authenticated route where you re-check permissions. Origin is the second lock, enforced by the browser. It answers: where does this code run? The answer should be “as far away from my main application as possible.” This is implemented by rendering the untrusted HTML inside a sandboxed iframe or by applying an extremely strict Content Security Policy (CSP) to the rendering page. The sandboxed iframe is the strongest option; it isolates the script’s execution context completely. Capability is the third lock, also enforced by the browser. It answers: what can the code do once it is running? Even in a sandbox, you must restrict what it can reach. This means disabling network access (sandbox attribute without allow-same-origin), blocking storage APIs, and preventing it from capturing user input beyond the iframe. The skill for a platform operator is implementing these boundaries with standard web features.
What this means in practice for your hosting
The operational difference is trusting the browser, not your upstream process. You implement these boundaries with concrete configurations. For identity, ensure your link-handling route performs a fresh authentication and authorization check on every request; do not rely on tokens passed in the URL string. For origin, the practical start is a Content-Security-Policy header. A policy like default-src 'none'; script-src 'self' is a good start, but for truly untrusted content, the sandboxed iframe is superior. You create an iframe with the sandbox attribute, which by default denies scripts the ability to navigate the top frame, access storage, and make network requests. The trade-off is that some interactive features, like embedded videos or external APIs, will break by design. That is the point. You are trading convenience for containment. For a client building an interactive prototype, this might mean their chart cannot fetch live data. That is an acceptable limitation when the alternative is a potential compromise of your entire user base.
What to expect next
I expect the volume and sophistication of AI-generated artifacts to increase steadily. Platform-side checks will also get better at catching obviously broken or benignly non-compliant output. But those upstream checks will always be playing catch-up. The sophistication of prompt injection and AI-generated exploits will advance at least as fast as the defensive models. The critical shift is that more code will be written by AI agents without any human review step in the loop. In that world, the identity-origin-capability model is not a best practice; it is the only viable practice. It becomes the essential infrastructure that allows you to unconditionally execute untrusted code, because the browser itself is now your primary security partner. Your architecture must reflect this reality, placing the browser’s enforcement mechanisms at the center of your security design, not as an afterthought.
Conclusion: Concrete next step for your platform
Audit how your current setup handles any user-provided or AI-generated HTML. Do you serve it from the same origin as your main application? Does it have full access to the DOM and network? If so, you are operating on an outdated trust model. The practical, immediate action is to implement a strict Content Security Policy. Start by adding a Content-Security-Policy header to the responses that serve these untrusted pages. A restrictive policy like default-src 'none'; style-src 'self' 'unsafe-inline' forces you to confront what resources the artifact truly needs and locks down everything else. For a deeper technical breakdown, consult the original post by Islam Gagiev (opens in new tab). The next step after CSP is evaluating a sandboxed iframe architecture. Build your platform assuming the code it serves is hostile, because with AI as the author, it is.