Intro: OpenAI models running inside your own AWS account
The Amazon Bedrock Managed Agents public preview puts OpenAI models behind the IAM roles, VPC endpoints, and CloudTrail trails you already run. According to Channy Yun’s October 5 roundup (opens in new tab), it is built on a customized version of OpenAI’s Agents API, engineered to be AWS-native and integrated with AWS resources. That sentence does more commercial work than the model list does. A team that could not get a new vendor through procurement in six months can often get a new endpoint approved in a week if it sits inside an account they already audit.
I have sat in the meeting where a client’s OpenAI key turned up as a Lambda environment variable, readable by anyone with lambda:GetFunctionConfiguration. That is not a hypothetical failure mode, it is the default outcome of letting developers move fast without a platform team in the loop. Moving the same workload onto a Bedrock-managed endpoint does not make it secure by itself, but it makes the security question answerable with tools the client’s own security team already knows how to query.
The preview is aimed at teams that are already on Bedrock for something else, or that have a platform group enforcing account structure across several business units. If you are a two-person shop running one agent, the pitch is thinner. You get governance you were not really using and a runtime you could approximate with a container and a cron job.
You can pick the execution environment. Self-hosted compute reuses a machine or container you already have, and AgentCore Runtime gives you managed runtime sessions with configurable storage in your AWS account. Which of those two you choose ends up mattering more for your data residency story than which model you point at.
Two things I could not confirm from the announcement and am not going to guess at. First, per-account quotas. Nothing in the roundup says how session concurrency is metered, and that number decides whether this works for a batch job or only for interactive use. Second, multi-region behavior. The roundup describes storage in your AWS account without specifying how a session pinned to one region behaves if the caller is not. I would want those in writing before designing around them.
Background: Bedrock’s roster and the lifecycle churn running alongside it
Bedrock started as a way to reach Claude models without leaving the AWS account, and for a long while that was the whole story. The catalog has since filled out with OpenAI, Anthropic, and xAI entries side by side, and once you are routing to several vendors from one endpoint, the next thing a platform team asks is where the agent loop itself lives. A managed agents layer is the predictable answer to that question.
The same announcement carries a service availability update, and the lifecycle vocabulary in it is worth pinning down because people mix the states up. Maintenance means the service still runs for existing customers but stops taking new signups, and in this roundup that starts October 29, 2026 for the Chime SDK SIP Media Application and WorkSpaces Secure Browser. Sunset means a shutdown date is now published: Managed Blockchain, DevOps Guru, and the Backint Agent for SAP ASE all end support in late September 2027, while the standalone Infrastructure Composer console goes on December 7, 2026. End of support is the terminal state, and as of September 29, 2026, Mechanical Turk sits there. All of this comes from the October 5 roundup (opens in new tab), which links out to the Product Lifecycle Changes page for the full lists.
I read the whole thing as a governance story that happens to have model names attached. Four new frontier models and a managed runtime make the headline, but the part with a deadline in it is the lifecycle table. Migration windows that close in 2027 have to be scoped in 2026, and a team that spends Q4 benchmarking GPT-6.1 Sol against Sonnet 5.5 is a team not reading the page that says their audit pipeline depends on a service with an end date.
If you operate anything on Managed Blockchain or DevOps Guru, the calendar is the news here, not the models.
What’s happening now: Amazon Bedrock Managed Agents powered by OpenAI
The preview is narrower than the name suggests, and that is mostly to its credit. AWS took OpenAI’s Agents API, customized it to be AWS-native, and wrapped it in the identity and permission model you already run. The agent code is OpenAI-shaped, but whatever executes it sits inside your account behind your existing IAM roles. You choose the execution environment: self-hosted compute (an existing development machine, a container, a compute environment you already pay for) or Amazon Bedrock AgentCore Runtime, which gives you managed runtime sessions with configurable storage in your own account. That single choice determines your blast radius, your latency profile, and your bill, so it is not a checkbox you tick during a proof of concept and forget.
Four models landed alongside it, and the vendor claims are worth separating from anything you can measure.
- OpenAI GPT-6.1 Sol, described as an upgrade to GPT-6 Sol aimed at agentic coding, computer use, and professional work. OpenAI says it approaches GPT-6 Astra on demanding evaluations at roughly one-fifth the cost. That comparison is OpenAI’s, repeated in the AWS roundup, not an independent benchmark.
- OpenAI GPT-6 Astra UltraFast mode, a premium speed tier where OpenAI claims up to 6x faster inference in the API at up to 300 tokens per second.
- Anthropic Claude Sonnet 5.5, positioned as a step up from Sonnet 5 for coding, with the pitch being well-scoped tasks inside a larger strategy: building and fixing features in one session, then verifying output against requirements.
- Grok 4.7, listed under SpaceXAI in the roundup, with better mixed-document handling, repo-scale coding that includes planning and error recovery, and browser-use agents for form fills and portal navigation.
Two infrastructure items in the same batch matter more to me than the model cards. The AWS Well-Architected Agent preview analyzes your environment and produces targeted recommendations on cost, security, performance, and resilience. “Contextual” here means the agent reads your account and your stated goals rather than handing you the generic Well-Architected checklist PDF you already ignored in 2021. Before wiring it up, decide what it can see, because an agent with broad read permissions is a new entry in your threat model.
The other is Aurora PostgreSQL now querying Apache Iceberg and Parquet directly, no ETL and no duplication, using the PostgreSQL clients you already have open. S3 Tables picked up full Iceberg V3 data type support in the same window: geometry, geography, unknown, and nanosecond timestamps, plus column defaults. If you have ever stored coordinates as "12.9716,77.5946" strings or millisecond epochs in a bigint column because the format did not support anything better, that workaround is now optional.
Amazon Bedrock Managed Agents: the questions I keep getting
Q: Do my prompts and data leave AWS?
This depends on the execution environment you pick, and that is the only part of the answer I can give with confidence. Self-hosted compute means the agent runs on a development machine, container, or compute environment you already control. AgentCore Runtime gives you managed runtime sessions with configurable storage inside your AWS account. What I cannot tell you from the roundup is where inference physically executes. The Agents API is described as a customized version of OpenAI’s API, engineered to be AWS-native and integrated with AWS resources, but “AWS-native” describes the integration surface, not necessarily the silicon. I would get that in writing from your account team before an agent touches customer data. My working assumption, and I am flagging it as an assumption, is that session state and storage stay in your account while the model call crosses whatever boundary OpenAI’s serving stack requires. Verify it rather than trusting my reading.
Q: Does this replace AgentCore or Strands Agents?
No. AgentCore Runtime is one of the two execution choices for managed agents, so it is the substrate rather than a casualty, and Strands shows up separately in the same roundup. If you already have tooling on either, the managed layer sits alongside it. The practical question is whether you want two agent frameworks in one repo, and usually you do not.
Q: What does “roughly one-fifth the cost” mean on my invoice?
That figure is OpenAI’s own positioning for GPT-6.1 Sol against GPT-6 Astra on demanding evaluations. It is a benchmark comparison from the vendor selling the model. Your invoice is token pricing multiplied by tokens actually consumed, plus runtime session charges if you go managed. An agent that needs four attempts to land a task is more expensive than a cheaper model that lands it in one, so measure cost per completed task and not cost per token. That number is the only one your finance team will recognise.
Q: Can I reuse my existing IAM roles and VPC setup?
Partly. Identities, permissions, and governance controls carry over, which is the entire pitch of the preview. What needs rework is the permission boundary. An agent typically wants read access to services your existing bedrock:InvokeModel role never touched, so a tightly scoped role from a static workload becomes an exercise in adding permissions rather than reusing them. On networking, self-hosted compute keeps your VPC setup as it stands. With AgentCore Runtime you get the managed session but storage still lands in your account, so VPC endpoints, bucket policies, and log destinations remain your responsibility.
What it means in practice
Model cards are marketing documents. “Approaches GPT-6 Astra across demanding evaluations” tells me nothing until my own workload runs through it. The benchmark that matters is the one I build: pull sixty real support tickets, or sixty invoices, or sixty log lines, feed them to both models with the same prompt, and count how many outputs I would ship without editing. That takes an afternoon. It has stopped me from two model migrations that looked excellent in a comparison table.
The Iceberg and Parquet support is the item here most likely to save actual engineering time. If you run Aurora PostgreSQL and your reporting data sits in S3 as Parquet, you currently either copy it into a table or maintain a Glue job that copies it. Querying it directly through the same database connection means your existing dashboards, your existing psql session, and your existing scheduled reports keep working unchanged. I have maintained those Glue jobs. They fail on a schema change on a Friday, every time.
Kiro workflows are the part I would approach more slowly. If your pipeline already runs on GitHub Actions or CodeBuild, adding an agent workflow layer gives you a second place where build logic lives and a second place to debug when a release breaks at 2am. I would want to see it replace a step, not sit beside one. A workflow engine that runs alongside your CI is a workflow engine you will eventually forget to update.
On cost and lock-in, self-hosted compute keeps your billing, autoscaling, and key management exactly as they are. AgentCore Runtime handles session state, storage, and scaling, which is genuinely work you would otherwise write and maintain yourself. That is a real trade. The part to think through now is where session state lands, because moving it back out to your own compute later is not a copy and paste job.
Pick one workflow you already understand well, run it both ways, and compare what you spent in engineering hours rather than what the model card promised.
What to expect next, starting with the 2027 dates
The lifecycle updates dated September 29, 2026 are worth reading against your own account, because the dates on that page have teeth. Amazon Chime SDK SIP Media Application and Amazon WorkSpaces Secure Browser both move to maintenance, and no new customers can sign up from October 29, 2026. Existing deployments keep running. If you have a voice application built on Chime SDK SIP media, that is the kind of service where you find out about the change from a customer asking why the console no longer offers it.
The sunset calendar runs longer. Amazon Managed Blockchain reaches end of support on September 29, 2027, Amazon DevOps Guru on September 30, 2027, and the AWS Backint Agent for SAP ASE shares that window. AWS Infrastructure Composer’s standalone console closes earlier, on December 7, 2026, while the service itself continues inside the main console. Long runways are the reason these get ignored for eighteen months and then become an emergency. Anyone running a consortium ledger on Managed Blockchain has a migration project that does not get cheaper by waiting.
Mechanical Turk is the item I keep coming back to. It reached end of support on September 29, 2026, with no future date attached. Compare that to a service entering sunset, which at least gets a stated deadline you can plan around. What the Mechanical Turk status signals to me is that maintenance is the real warning shot. By the time AWS moves something to end of support, the decision has already been made.
My expectation, and this is an expectation rather than anything AWS has published, is that per-account quotas and multi-region behaviour will be what actually gates adoption of the Bedrock Managed Agents preview. Model access is easy to grant. Session storage that spans regions, or a quota that caps concurrent agent sessions at a number you cannot raise on a support ticket, is the thing that stops a production rollout. I would also want the documentation to state plainly whether session state is encrypted with a key you control, and what happens to it when a run finishes.
Until that is written down, I would treat the preview as exactly that: a place to build one agent, learn the session model, and keep the production one on infrastructure you already understand.
Next step: run one agent inside your own account before you decide anything
The preview is free to explore, which is the whole point of exploring it. Pick a single low-risk agent, something that summarises a CloudWatch alarm or drafts a reply to a support ticket, and give it an IAM role that can read exactly one resource and nothing else. No wildcard policies on day one. Run it on AgentCore Runtime so you actually experience the managed session model rather than guessing at it from documentation, and watch where the session data lands. If the storage location and the retention behaviour feel wrong, you have learned that for the cost of an afternoon instead of a migration project.
The IAM role is worth scoping by hand rather than through the console wizard. Write the policy as JSON, set a permissions boundary, and check what the agent can reach with a deliberate test call it should fail. Agents that inherit broad permissions tend to accumulate capability quietly, and you will not notice until one of them touches something it was never meant to.
Read the model card for the model you intend to run, not the one with the strongest comparison table. GPT-6.1 Sol, GPT-6 Astra in UltraFast mode, Claude Sonnet 5.5, and Grok 4.7 are marketed against different workloads, and the phrasing in each card tells you which one the vendor expects you to reach for. Pricing per token does not describe what a run costs when retries and tool calls are involved. Run your own prompt set through two of them and count the tokens that actually moved.
Then do the boring audit. Open the AWS Product Lifecycle Changes page, and while you are there, check AWS Health in your own account for anything sitting in maintenance or carrying a sunset date you have not actioned. Subscribe to the Health Dashboard notifications if you have not already, because that is how the Chime SDK and WorkSpaces Secure Browser dates will reach you rather than a blog post eighteen months later. The date to circle right now is October 29, 2026, when new customers lose access to those two services. If you are already a customer, you have runway. If you were planning to become one, that plan just changed.