Welcome to episode 364 of The Cloud Pod, where the forecast is always cloudy! Justin and Matt are in the studio this week to bring you all the latest in cloud and AI news, including (surprise) astronomical AWS bills, Kimi K3 and what it means for Enterprise AI, and lots of security news! All that and so much more, so let’s get started!
Titles we almost went with this week
- 👷 Cache Rules Everything Around NFS Now
- 🪣 Lambda Says BYOB, Bring Your Own Bucket
- 🪫 Henrico’s Power Struggle: Data Centers 37, Schools 0
- 🏃 Cloud Run Fails Over Faster Than Your Excuses
- 💭 570 Patches, One Registry Hive Nightmare
- 🔍 GuardDuty Gets a Detective Agent, No Trench Coat Required
- 🌛Kimi K3 Aims to Moonwalk Past Opus 4.8
- 🍎 Terraform Gets Policy Muscle, Ditches the Rego Diet
- ☁️ CloudWatch Watches Your AI Coders Code
- 💡 Watt A Way To Treat A School District
- 💰 Your Cloud Bill… 1 BILLION DOLLARS
- 💸 Not a way I want to wake up rogue cloud bills
- ⚰️ Skype is EOL … wait I thought I died 3 times already
- 🦸 Our newest superhero CODEMENDER!!
A big thanks to this week’s sponsors:
We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info.
General News
00:49 Amazon fixing bug that billed some AWS customers billions of dollars
- A bug in the AWS billing computation subsystem generated inflated billing estimates for some customers, with one Reddit user reporting a quoted estimate near 2.5 billion dollars for a single month. In contrast, others saw figures ranging from millions to hundreds of millions.
- The issue began late Thursday, and an initial rollback attempt on Friday morning failed to resolve it, suggesting the root cause was more complex than a recent configuration change.
- Amazon confirmed the billing estimates do not reflect actual usage or charges, meaning affected customers will not be responsible for the inflated amounts shown in the console.
- Amazon has not disclosed whether any accounts were suspended or paused due to the billing errors, leaving open questions about operational impact during the incident.
- The event highlights the importance of billing system reliability for cloud providers, since inaccurate estimates at this scale can cause confusion and concern even when the underlying charges are not real.
01:29 📢 Justin – “Amazon doesn’t bill you in the middle of the month, so it’s a pretty low risk that you were gonna get billed or invoice directly on that date, unless you happen to already be overdue on a payment and you were happening to update your credit card at the same time. I don’t think that’s really a big risk for this particular scenario.”
05:04 County With 37 Data Centers Asks Schools to ‘Conserve Electricity’
- Listener note: Paywall article
- Henrico County, Virginia, home to 37 data centers, with 17 more planned, is asking county employees and schools to conserve electricity after a 25 percent rate increase set to begin July 1, adding an estimated 5 million dollars in costs for the next fiscal year.
- The situation highlights a direct tension between data center growth and local infrastructure costs, with residents and government facilities absorbing higher utility rates likely tied to the power demands of nearby facilities.
- Proposed expansion includes converting Civil War battlefield land into data center space, raising questions about land use and community pushback in addition to energy concerns.
- This case illustrates a broader pattern playing out in data center hub regions nationwide, where rapid buildout strains local power grids and shifts cost burdens onto residents and public institutions rather than solely the operators.
- Worth discussing on the podcast: how utilities allocate rate increases across commercial and residential users, and whether data center operators like Meta and others in the county are contributing to infrastructure upgrades or offsetting costs for the community.
08:20 📢 Justin – “Don’t take power away from school kids. And… it sounds like in a lot of the newer municipalities where they’re agreeing to put these data centers in, they’re saying we’re not pushing rate increases down onto the general population; that if rate increases are required because you’re using so much power, you’re gonna pay for it, which I think is the right way to handle that.”
AI Is Going Great – or How ML Makes Money
09:54 GPT-Red: Unlocking Self-Improvement for Robustness
- OpenAI trained GPT-Red, an internal-only automated red-teaming model used to find prompt injection vulnerabilities and generate adversarial training data at the compute scale of some of its largest post-training runs.
- GPT-Red uses self-play reinforcement learning against a population of defender LLMs, with GPT-Red rewarded for successful attacks and defenders rewarded for resisting them, forcing progressively stronger and more diverse attack discovery.
- Incorporating GPT-Red into training produced GPT-5.6 Sol, which shows 6x fewer failures on the hardest direct prompt injection benchmark versus the production model from four months earlier, and fails on only 0.05 percent of GPT-Red’s direct prompt injection attempts.
- In generalization tests, GPT-Red achieved an 84 percent attack success rate on novel indirect prompt injection scenarios against GPT-5.1, compared to 13 percent for human red-teamers on the same tasks.
- In a real-world test against an AI-powered vending machine agent (similar to Anthropic’s Project Vend), GPT-Red successfully changed item pricing, created a fraudulent listing, and canceled another customer’s order, with the vulnerabilities disclosed and safeguards now being tested.
- OpenAI reports that capability evaluations and over-refusal tests show robustness gains came from better resistance to malicious instructions rather than the model becoming more restrictive or less capable overall.
11:15 📢 Justin – “The ability to attack and attack from multiple vectors and chain attacks is only increasing dramatically at this point.”
12:19 OpenAI’s first branded hardware is… a light-up keyboard?
- OpenAI released its first branded hardware, the $230 Codex Micro, a collaboration with Work Louder built on their existing Creator Micro keyboard line, rather than a fully in-house design.
- The keyboard’s key feature is six frosted, color-coded keys that provide status updates on up to six concurrent Codex agent threads: white for idle, blue for processing, green for completed, amber for needing human input, and red for errors.
- Six additional programmable buttons handle common Codex actions like accepting or rejecting changes and branching threads, plus a push-to-talk button for audio prompts; users can remap these and access five additional customizable layers for general shortcuts via 32 included keycaps.
- The device addresses a workflow problem for developers running multiple AI coding agents simultaneously, offering at-a-glance monitoring as an alternative to keeping several browser tabs or a laptop open to track agent status.
- This is a desktop-focused accessory that complements rather than replaces mobile monitoring options like the ChatGPT app, and it arrives alongside ongoing reports of OpenAI developing a separate screenless AI companion speaker for release in coming years.
- We don’t get this. Any listeners out there planning on grabbing this? Let us know.
15:55 Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8
- Moonshot AI’s upcoming Kimi K3 model is reported to perform on par with or exceed Anthropic’s Opus 4.8, according to sources cited by the Financial Times. It’s expected to be the largest open-weight AI model out of China, with parameters ranging between 2 and 3 trillion.
- The predecessor, Kimi K2, already ranks competitively on open-source benchmarks, and K3 aims to further narrow the performance gap with closed-source frontier models from OpenAI and Anthropic.
- Moonshot is reportedly raising a new funding round at a $31.5 billion valuation, up from $20 billion in May when it raised $2 billion, reflecting continued investor interest in open-source AI development.
- This release comes as enterprise leaders debate the cost and data privacy tradeoffs of closed-source AI subscriptions, with some executives recommending open-source alternatives like Moonshot, DeepSeek, or Z.ai for organizations wanting to train and control their own models.
- For cloud and infrastructure teams, a competitive open-weight model at this parameter scale could shift self-hosting economics, giving enterprises more leverage in negotiations with closed-source providers or a viable path to bring model training and inference in-house.
17:00 📢 Justin – “The overall stock market has not been favorable to this this week because again, there’s a lot of companies investing a lot of capital, and these cheaper models put that business model at risk. And so the market is appropriately reacting this week. But I’m definitely excited to get my hands on Kimmy K3.”
18:51 Kimi K3 Tech Blog: Open Frontier Intelligence
- Kimi K3 is a 2.8-trillion-parameter open-weight model, described as the first open model at 3T-class scale, built with a 1-million-token context window and native vision support.
- Full model weights are scheduled for release by July 27, 2026, with the model already usable via Kimi.com, Kimi Work, Kimi Code, and the Kimi API.
- Architecture relies on two new components, Kimi Delta Attention and Attention Residuals, plus a Stable LatentMoE setup activating only 16 of 896 experts. The company reports roughly a 2.5x improvement in scaling efficiency compared to its prior K2 model.
- Benchmarks show K3 trailing the top proprietary models, Claude Fable 5 and GPT 5.6 Sol, but consistently ahead of other open and proprietary models tested across coding, knowledge work, and agentic tasks, including DeepSWE, Terminal-Bench 2.1, and BrowseComp evaluations.
- Notable case studies include building a GPU compiler called MiniTriton from scratch that matches or beats Triton on some workloads, designing a functional chip in a 48-hour autonomous run, and completing a two-week astrophysics research task in about two hours.
- API pricing is set at 0.30 dollars per million tokens for cache-hit input, 3.00 dollars for cache-miss input, and 15.00 dollars for output, with a reported cache hit rate above 90 percent for coding workloads via Mooncake’s disaggregated inference architecture.
- The company recommends deployment on supernode configurations with 64 or more accelerators for optimal inference efficiency.
- Documented limitations include instability when thinking history isn’t properly preserved across sessions, a tendency toward excessive proactive decision-making on ambiguous tasks, and an acknowledged gap in overall user experience compared to Claude Fable 5 and GPT 5.6 Sol
20:24 Introducing the ChatGPT for small business program
- OpenAI launched the ChatGPT for small business program, bundling virtual training webinars, in-person AI academies, guides, and curated partner integrations from Dropbox, Shopify, Intuit, Slack, Atlassian, and Wix.
- ChatGPT Work, OpenAI’s multi-step task agent, is now available to small businesses and runs on GPT-5.6, positioned as OpenAI’s most advanced model available across all business subscription tiers.
- From last year’s Small Business AI Jam events, OpenAI reports 78 percent of participants built a functional AI workflow in a single day, and 42 percent saved more than five hours per week using AI tools.
- Use cases highlighted include converting voice notes into Slack messages, generating real-time market/competitor tracking sites, evaluating inventory for product or marketing ideas, and building training presentations from customer review data.
- The program targets a segment often lacking dedicated IT or automation resources, framing agentic AI as a way to offload tasks like marketing, accounting, and operations that would otherwise require outsourcing or additional hires.
Security
22:08 Windows 0-day drops the same day Microsoft releases record number of patches
- Microsoft released 570 security patches, a record volume for a single update cycle, and a zero-day exploit surfaced the same day affecting the Windows User Profile Service.
- The exploit, called HiveLegacy, allows a low-privilege account to modify an administrator account’s classes registry hive, which controls file association behavior in Windows Explorer.
- Exploitation requires the attacker to know credentials for one account and the username of a second account on the same machine, limiting but not eliminating practical risk.
- The researcher, using the pseudonym NightmareEclypse, has published nine such exploits and has stated dissatisfaction with Microsoft’s handling of vulnerability disclosures, raising questions about coordinated disclosure practices.
- The volume of patches combined with an active zero-day highlights ongoing challenges in patch management and prioritization for IT teams managing Windows environments at scale.
22:40 📢 Justin – “570 security patches is a LOT of security patches, and I can definitely thank AI for all of those patches.”
Cloud Tools
27:07 1Password for Claude: Give Claude access without giving up your credentials
- 1Password for Claude lets the AI agent complete browser logins and tasks without ever seeing the actual password or one-time passcode; credentials are injected directly into the page at runtime while 1Password remains the source of truth.
- Access is scoped per-task and requires explicit user approval via biometric confirmation each time Claude needs a credential; after autofill, 1Password verifies secrets weren’t exposed on the page and clears values if submission fails.
- Agentic Mode addresses a separate risk: when a browser agent takes control of a browser with 1Password installed, the extension locks down automatically, hiding the UI and restricting the agent to only pre-approved logins, leaving the rest of the vault inaccessible.
- The integration is available now for Mac across business, family, and individual plans, and fits into 1Password’s broader strategy of building a trusted access layer for AI agents, including similar MCP server integrations for OpenAI Codex and Kiro.
- This addresses a practical security gap as agents move from advisory roles to taking real actions like purchases and account changes, framing AI agents as a new identity class requiring the same governed, runtime-scoped access model as human or machine identities.
27:24 📢 Justin – “It exists – I can’t make it work.”
28:39 Introducing tfpolicy: A declarative policy workflow built for Terraform
- HashiCorp launched tfpolicy in public beta on HCP Terraform, a declarative policy-as-code framework using HCL instead of separate languages like Sentinel or OPA rego, letting platform teams write governance rules in the same syntax used for infrastructure definitions.
- A key new capability is relationship-aware policy evaluation, allowing rules to span multiple connected resources, for example requiring every IAM role to have at least one attached policy rather than checking resources in isolation.
- The framework supports data source lookups during policy evaluation, so policies can reference external context like approved AMI lists or organizational inventories rather than relying solely on what’s defined in the Terraform configuration.
- Tfpolicy adds controls to block unapproved provider and module downloads before use, addressing supply chain risk by enforcing that dependencies come from approved private registries.
- Policies can also be evaluated post-deployment, checking provider-computed values like generated ARNs against organizational standards, which addresses gaps where plan-time checks alone are insufficient.
- HashiCorp is providing an agent skill on GitHub to help teams author and test tfpolicy files or migrate existing Sentinel policies, easing adoption for current HCP Terraform customers.
29:58 📢 Justin – “I suspect that Sentinel is going to go away – or at least be heavily deprecated in favor of this method.”
AWS
32:27 AWS Lambda announces self-managed code storage
- Lambda now lets functions and layers reference code directly from customer-owned S3 buckets instead of copying deployment packages into Lambda-managed storage, removing the 75GB per-Region storage cap for those using this mode.
- Teams with many functions or large layers no longer need to file support tickets to raise storage quotas.
- Skipping the internal copy step also reduces function activation time after creates and updates, which benefits customers with large deployment packages or frequent deployment cycles.
- Setup requires setting S3ObjectStorageMode to REFERENCE via CLI, CloudFormation, SAM, or SDKs, plus granting the Lambda service principal s3:GetObject and s3:GetObjectVersion permissions on the source bucket. Console-based updates are also supported for existing functions.
- No additional Lambda fees apply for self-managed storage; customers pay standard S3 storage rates and cross-Region data transfer costs where applicable, making this a cost-neutral change for most workloads.
- AWS also raised the default Lambda-managed code storage limit from 75GB to 300GB per Region per account, benefiting customers who don’t migrate to self-managed storage.
- The feature is available now across all commercial AWS Regions.
33:50 📢 Matt – “I’ve done a lot of development, when Terraform first came out, I would have my Lambda built into my Terraform. I would just do a Terraform Ply every time, which would just zip up the folder and shove it into Lambda. So I’ve never hit that limit.”
35:01 Amazon MQ now supports configurable storage for RabbitMQ brokers
- Amazon MQ now lets customers configure EBS storage size independently of instance type for RabbitMQ brokers, addressing a long-standing limitation where storage was tied to compute sizing.
- The feature is limited to RabbitMQ M7g brokers on version 4.2 or later, and only supports cluster deployments, so single-instance broker users won’t have access to this option.
- Storage can be adjusted in 5 GB increments up to the maximum allowed for the instance size, configurable via AWS Console, CloudFormation, CLI, or CDK, though changes only take effect after the next broker reboot.
- This decoupling allows customers to right-size costs for messaging workloads with high storage needs but modest compute requirements, or vice versa, avoiding the need to overprovision instance size just to get more disk space.
- Pricing follows standard Amazon MQ storage rates based on disk size, with no additional fees for the configurability itself, and the feature is available in all commercial regions where Amazon MQ for RabbitMQ is offered.
35:23 📢 Justin – “We talked about RabbitMQ last week, and I said, yeah, I don’t care about that. And apparently Amazon still does enough care and cares enough to still build features for it. So there you go.”
36:19 Amazon CloudWatch Logs announces intelligent tiering for storage
- CloudWatch Logs now automatically tiers data into Standard, Infrequent Access, and Archive Instant Access based on usage, removing the need to manually filter or export logs to cheaper storage elsewhere.
- Data shifts to Infrequent Access after 30 days without access and to Archive Instant Access after 90 days, with automatic promotion back to Standard for 30 days when older logs are queried again.
- Query experience remains consistent across all tiers, letting teams keep verbose, high-volume logs in CloudWatch long-term without switching tools or maintaining separate storage systems.
- Consolidating logs in one place simplifies operations and could reduce Mean Time to Resolution by keeping all data queryable and alertable from a single service.
- Available in all AWS commercial regions except Middle East (Bahrain) and Middle East (UAE); can be enabled account-wide via the console, SDKs, or CLI. Pricing details are on the CloudWatch pricing page.
37:18 Amazon Cognito now supports importing users with password hashes
- Amazon Cognito now allows password hashes to be included in CSV user imports, letting migrated users sign in immediately with existing credentials instead of being forced into a password reset on first login.
- Supported hashing algorithms include bcrypt, scrypt, Argon2id, and PBKDF2 with SHA-256, covering most common formats used by legacy identity systems and custom auth implementations.
- Imported hashes receive an additional layer of cryptographic protection before being stored in Cognito, addressing a key security concern for teams migrating user directories.
- This directly targets the migration pain point of moving off a legacy IdP or homegrown auth system, reducing user friction and support overhead during cutover.
- Available now in all AWS regions where Cognito operates, accessible via the Console, CLI, or SDKs, with no additional pricing beyond standard Cognito user pool costs.
38:09 📢 Justin – “Thank God. This was such an annoyance.”
39:38 AWS Control Tower Account Factory for Terraform now re-applies customizations when accounts move between OUs
- AWS Control Tower Account Factory for Terraform (AFT) now automatically re-applies account customizations when accounts move between Organizational Units, eliminating the manual re-triggering step that previously created operational overhead and configuration drift risk.
- Enable the feature by setting aft_customization_triggers equal to account_move in your AFT configuration; the re-application process skips bootstrap and provisioning phases, running only global and account-level customizations for faster execution.
- Teams retain granular control through account_skip_customization_triggers, which allows specific accounts to opt out of the automated re-application behavior when needed.
- This update is particularly relevant for organizations enforcing compliance or security baselines tied to OU membership, ensuring accounts stay aligned with policy requirements immediately after an OU move rather than during the next scheduled sync.
- The release also includes secondary improvements: custom Terraform Cloud and Enterprise workspace naming variables, tighter access controls on the AFT logging bucket, and improved scaling for large-scale AWS Enterprise Support enrollment. Available now in all regions where AFT is offered, with no additional cost beyond standard Control Tower and underlying resource usage.
40:03 📢 Justin – “Thank you. This was dumb.”
41:06 AWS Sustainability service now includes water withdrawals data
- AWS Sustainability now adds water withdrawals data alongside existing carbon emissions metrics, giving customers a fuller picture of the environmental footprint tied to their workloads.
- Data is broken down by AWS Region, service, and account, and is reported annually through both the console and API, allowing teams to integrate it into existing reporting workflows or dashboards.
- The feature is free in all Regions where AWS Sustainability is available, removing cost as a barrier to adoption for ESG and sustainability reporting teams.
- Lower withdrawal volumes reflect data center efficiency improvements, giving customers a way to track AWS infrastructure efficiency gains over time as part of their own sustainability disclosures.
- This addition is relevant for organizations facing increasing regulatory or investor pressure to report water usage as part of broader environmental, social, and governance (ESG) commitments, not just carbon metrics.
42:50 Amazon S3 removes 30-day minimum for transitions to S3 Standard-IA and
- AWS eliminated the 30-day minimum retention requirement for transitioning S3 objects to Standard-IA and One Zone-IA, allowing lifecycle rules to move data as soon as 0 days after creation.
- This change directly benefits workloads where data cools quickly, such as backups, log analytics, and compliance archives, letting customers capture up to 40% storage cost savings without the previous waiting period.
- Previously, customers had to keep data in S3 Standard for 30 days before transitioning, even if the data was rarely accessed after creation, so this removes an artificial cost inefficiency for short-lived hot data.
- Implementation is straightforward through updated S3 Lifecycle rules via console, CLI, or SDK, and the feature is available in all regions where these storage classes already exist, requiring no migration or architectural changes.
- Worth discussing how this affects cost optimization strategies for customers with predictable data access patterns, particularly those generating high volumes of logs or backups that are rarely read after initial creation.
43:48 📢 Matt – “It’s a great quality of life improvement. I’ve definitely inadvertently set things to these and then deleted them or moved them and then got hit with a fee… I’ll take the win and move on in life.”
45:20 Amazon CloudWatch announces coding agent insights
- CloudWatch coding agent insights gives engineering leaders visibility into AI coding tool usage and ROI, integrating with Claude apps gateway for AWS to pull telemetry from Claude Code without extra instrumentation; Codex and GitHub Copilot are also supported.
- The feature is built on OpenTelemetry metrics and surfaces them alongside existing CloudWatch operational data, letting teams correlate agent adoption with commit throughput, pull request velocity, and cost-to-output ratios by model.
- Practical use cases include setting proactive token billing alerts, tracking spend trends by department, and identifying which teams would benefit from expanded coding agent access.
- Availability spans all AWS commercial regions except Middle East (UAE), Middle East (Bahrain), and Israel (Tel Aviv); setup requires configuring the Claude apps gateway to emit telemetry to CloudWatch per the setup guide.
- Pricing follows standard CloudWatch OpenTelemetry metric ingestion rates, so costs scale with metric volume rather than a flat fee; check the CloudWatch metrics pricing page for specifics.
46:46 📢 Justin – “The token maxxing era was glorious for moments – and now it’s over.”
46:56 Selectively log network activity events by identity in AWS CloudTrail
- CloudTrail now supports IAM identity-based filtering for network activity events tied to VPC endpoints, letting teams log only relevant traffic instead of every API call passing through a PrivateLink connection.
- Practical use case: configure selectors to capture VpceAccessDenied events only from identities outside a trusted allowlist, which helps flag potential data exfiltration attempts while suppressing noise from known, approved roles.
- This supports data perimeter strategies by combining UserIdentity conditions with existing selector fields like eventName or vpcEndpointId, giving security teams granular control over what gets recorded.
- Reduces both log volume and CloudTrail costs since routine traffic from trusted principals no longer needs to be logged, while still preserving visibility into anomalous or unauthorized access patterns.
- Available now via Console, CLI, and SDKs in all regions where CloudTrail network activity events are supported, with no new service to provision, just updated advanced event selectors.
47:41 📢 Matt – “It’s great that you can actually start to select what you need. There was so much noise in there, and finding stuff was like a needle in the haystack, even once you followed all their guides and pumped it to Athena and then to your S3. You had Athena, and we went down that whole path and then tried to search it, but still finding the denial in there…”
49:16 Introducing the Amazon GuardDuty investigation agent: on-demand AI-powered threat assessment
- GuardDuty investigation agent, now in public preview, automates security finding correlation and investigation, cutting analysis time from hours to minutes by providing risk levels, confidence scores, MITRE ATT&CK mapping, and prioritized remediation steps.
- Investigations can be scoped to a single finding, a specific account, or an entire organization, and can be triggered via console, CLI, API, or natural language prompts up to 2,048 characters describing areas of concern.
- The agent integrates with the AWS MCP server, allowing teams to invoke investigations through natural language via tools like Claude or Kiro, and fits into existing pipelines (for example, EventBridge to SIEM) so Lambda functions can enrich raw findings with structured assessments before routing to incident response queues.
- This is distinct from AWS Security Incident Response, which pairs AI agents with human engineers for active incidents; the GuardDuty investigation agent is for on-demand assessment rather than incident coordination.
- Available at no charge during public preview in 10 regions including us-east-1, us-west-2, and eu-west-1, with usage capped at 10 investigations per account per day and 100 total during the preview period; investigation completion times run 2-5 minutes for account-level scope and 10-12 minutes for individual finding investigations.
50:09 📢 Matt – “This sounds pretty cool. I’d be interested in what it’s going to cost in the long term because it always worries me with that, but I think it can have a lot of value.”
51:15 Amazon SES introduces pricing plans
- Amazon SES now offers three bundled pricing tiers, Essentials, Pro, and Enterprise, replacing the previous model where deliverability features were purchased individually as add-ons.
- Each tier builds on the last: Essentials covers deliverability insights, Pro adds managed dedicated IPs, email validation, and inbox placement visibility, and Enterprise includes multi-region resilience, workload-level reputation isolation, and annual deliverability assessments.
- The bundling approach is aimed at simplifying procurement for customers who previously had to evaluate and purchase capabilities separately, with AWS noting the plans are discounted compared to à-la-carte pricing.
- Available in all SES regions except Middle East (UAE) and Middle East (Bahrain); customers can select a plan directly from the SES console pricing plan section.
- Worth discussing how this shift reflects a broader trend of AWS packaging services into tiered plans rather than pure consumption pricing, similar to enterprise software licensing models.
GCP
54:09 NotebookLM is now Gemini Notebook
- NotebookLM has been rebranded as Gemini Notebook, reflecting its expansion beyond a standalone research tool into deeper integration with the Gemini app and Google Search, following adoption by over 30 million users and 600,000 organizations since its 2023 launch as Project Tailwind.
- A key technical update gives each notebook a secure cloud computer, enabling native code execution for more complex data analysis grounded directly in user-provided sources. This is currently available to Google AI Ultra users and Workspace customers with AI Ultra or AI Expanded Access, with a rollout to Pro users on the web planned in the coming weeks.
- Cross-app syncing now connects the Gemini app and standalone Gemini Notebook, and Google plans to bring notebooks into AI Mode in Search, expanding where and how users can access research tools within the broader Google ecosystem.
- Use cases span business onboarding materials and student study aids, such as converting notes into audio or video summaries, indicating broad applicability across professional and educational contexts.
- No specific new pricing was announced beyond existing Google AI Ultra and Workspace tiers; access to the new cloud-computer feature is currently tied to those subscription levels, with broader availability to Pro users expected soon.
54:41 📢 Justin – “This is just this is a no-brainer. Like, yes, you take Notebook, you turn that into more of a CoWork type solution on top of Gemini Enterprise, you rebrand it, and now everyone thinks it’s all one product; which Google desperately needs brand marketing help on this stuff.”
56:14 Cloud Run multi-region services enhanced for high availability
- Cloud Run now supports automated failover for multi-region services with two new capabilities: readiness probes for instance-level health checks and service health aggregation across regions, exposed via serverless NEGs.
- When paired with a global external application load balancer, traffic automatically shifts away from unhealthy regions within seconds, removing the need for manual incident response during regional outages.
- Two deployment paths are supported: a global external application load balancer for public-facing apps, and a cross-regional internal application load balancer for private VPC traffic.
- Works best in active-active configurations with read/write-heavy workloads that synchronize data across regions; teams still need to handle redundancy at the database layer separately, using services like Spanner, Firestore, Cloud SQL, or Cloud Storage in multi-region configurations.
- Available now in all Cloud Run regions at no additional feature cost, customers only pay standard CPU and memory charges for running the readiness probes. Documentation available here.
58:07 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- Google released three new Flash models: 3.6 Flash, 3.5 Flash-Lite, and a specialized 3.5 Flash Cyber for security use cases, all targeting improved efficiency for production AI agents.
- 3.6 Flash uses 17% fewer output tokens than 3.5 Flash while improving coding, knowledge work, and multimodal benchmarks like OSWorld-Verified (83.0% vs 78.4%) and MLE Bench (63.9% vs 49.7%).
- Pricing is notably lower with 3.6 Flash at $1.50 per 1M input tokens and $7.50 per 1M output tokens, reducing overall cost per agentic task compared to 3.5 Flash. 3.5 Flash-Lite is priced at $0.30 per 1M input tokens and $2.50 per 1M output tokens, running at 350 output tokens per second, making it suited for high-throughput workloads like document processing and agentic search.
- 3.5 Flash-Lite reportedly outperforms the larger 3 Flash model on some benchmarks, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%), giving developers a faster and cheaper alternative for certain coding and agentic tasks.
- Both models support configurable thinking levels, letting developers balance latency and cost against reasoning depth.
- 3.5 Flash Cyber is a specialized model fine-tuned for vulnerability detection and patching, deployed through Google’s CodeMender agent using multiple coordinated model instances to generate consolidated security reports. Access is restricted to governments and trusted partners via a limited pilot program due to the dual-use risk of cybersecurity-focused AI.
- Both 3.6 Flash and 3.5 Flash-Lite are available now through Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app, with Flash-Lite also rolling out in Google Search.
- Google also confirmed 3.5 Pro is in partner testing and disclosed that pre-training has begun for Gemini 4.
59:19 📢 Justin – “If you’re into the flash model… you don’t need to use the Gemini Pro models very often. Although, like image generation, I use the more pro image models typically, because they listen to me better than the flash ones do. But nice to see these are getting updated once again.”
1:00:00 Now in preview: Find and fix software vulnerabilities with CodeMender
- CodeMender, born from Google DeepMind research, moves into preview as a managed AI agent that scans, verifies, and remediates code vulnerabilities, available via Gemini Enterprise Agent Platform or as part of AI Threat Defense.
- The agent goes beyond static analysis by building and running proof-of-concept exploits in a customer-managed sandbox to confirm a vulnerability is actually exploitable, reducing false positives and alert fatigue before generating a fix.
- Remediation is delivered as a code diff for developer review, with an LLM-as-a-judge step checking that patches don’t break existing functionality; developers retain approval control before anything is committed to the repository.
- Supports common languages including C/C++, Go, Java, Python, Ruby, Rust, and TypeScript, and integrates into CI/CD pipelines, VS Code, Antigravity, or a CLI client for local development workflows.
- Follows a multi-model approach, letting teams pick models based on cost, speed, or scanning depth, with third-party frontier model support planned later this year; a specialized version with Gemini 3.5 Flash Cyber is limited to select government and trusted partner access initially.
- Integration with AI Threat Defense uses Wiz to orchestrate the workflow, calling CodeMender to scan code, enrich findings via the Wiz Security Graph, and trigger Wiz Red Agent for AI pentesting to prioritize the highest-risk issues; early customer quotes come from Salesforce, Robinhood, and Palo Alto Networks.
Azure
1:01:49 Microsoft expands Azure AI and HPC infrastructure with AMD
- Microsoft is expanding Azure infrastructure with three new AMD-powered VM families: HDv2 for data processing, HXv2 for electronic design automation, and ND MI455X v7 for AI inference, all built on AMD’s Helios platform and next-gen EPYC CPUs.
- HDv2 targets CPU-heavy AI workloads like data prep and agent coordination, offering nearly 500 physical 6th Gen EPYC cores, 4TB RAM, 32TB local NVMe storage, and 400 Gb networking, addressing the CPU bottleneck that can starve GPU accelerators of data.
- HXv2 builds on the 2023 HX series with 3D V-cache technology, now featuring 176 EPYC cores at over 5 GHz, 50% more cache per core, up to 4TB RAM, and 800 Gb InfiniBand, aimed at chip design firms running RTL simulation and broader HPC workloads like scientific simulation and MPI-based applications.
- ND MI455X v7 is positioned for production-scale AI inference, reasoning, and agentic workloads, using AMD’s Helios rackscale architecture, giving customers another inference option alongside Microsoft’s own custom silicon.
- The announcement reinforces Microsoft’s multi-vendor silicon strategy, pairing AMD hardware with in-house chips to offer customers workload-specific compute choices rather than a one-size-fits-all approach; no pricing details were disclosed, and availability timing wasn’t specified in the announcement.
1:02:17 📢 Matt – “Their naming convention makes sense if you understand and you have the translator for it.”
1:03:01 Reminder: Skype for Business 2015 and 2019 ESU Program Ends in
- Microsoft confirmed there will be no further extension of the Skype for Business 2015/2019 Extended Security Update program beyond October 2026, closing out the “Period 2” ESU that followed an earlier one-time extension.
- Organizations still running Skype for Business 2015 or 2019 in production will receive no further security updates after October 2026, creating a hard deadline for migration planning.
- Microsoft is steering customers toward an in-place upgrade path from Skype for Business Server 2019 CU8 to Skype for Business Server Subscription Edition (SE), which it describes as low risk since it is not a significant technological change.
- Skype for Business Server 2015 users have a more involved migration path since mainstream support for that version already ended, requiring a different approach than the SE in-place upgrade.
- This is a relevant reminder for IT admins and podcast listeners managing on-premises unified communications infrastructure, as missing the October 2026 cutoff means running unsupported, unpatched software.
1:03:39 📢 Justin – “I thought it died like three times already.”
Emerging Clouds
1:04:48 Upcoming GPU Pricing Updates
- DigitalOcean is raising on-demand pricing for NVIDIA and AMD GPU droplets effective August 1, 2026, citing strong demand for GPU capacity; existing customers must destroy droplets before that date if they want to avoid the new rates.
- Billing changes will apply to any active workloads running on or after August 1, 2026, with the updated charges appearing on the September 1, 2026 bill, giving customers roughly a month’s lag before seeing the impact.
- Reserved 12-month pricing is also increasing, but customers currently under contract keep their locked-in rate until renewal, which incentivizes existing customers to consider extending contracts before the change takes effect.
- This signals a broader trend of GPU cloud providers adjusting prices upward as demand for AI training and inference capacity continues to outpace supply, a pattern worth watching across other providers.
- For teams with predictable workloads, this reinforces the value of reserved capacity versus on-demand pricing, and businesses should evaluate their usage patterns now to lock in current rates before the August deadline.
1:05:25 📢 Justin – “This is just the reality of the continuing pressure on the compute market… unfortunately it’s happening to DigitalOcean, but it’s also happening everywhere.”
Closing
And that is the week in the cloud! Visit our website, the home of the Cloud Pod, where you can join our newsletter, Slack team, send feedback, or ask questions at theCloudPod.net or tweet at us with the hashtag #theCloudPod

Leave a Reply