r/aws 5h ago

discussion After last week's Cost Management display bug, how do you actually validate your dashboard against source-of-truth?

0 Upvotes

When your Cost Management dashboard shows you 7 trillion dollars, the actual problem is not that number. The actual problem is that every other number the same dashboard has ever shown you just lost credibility.

Last week Cost Management surfaced anomaly figures in the trillions and billions across a bunch of accounts. AWS Service Health confirmed a metering pipeline display bug. The real bill for most of those accounts was pocket change. Bug patched inside a few days.

The part that stuck with me is not the bug. It is the aftershock. If the dashboard can render 7 trillion, and I only know it is wrong because the number is physically impossible, then what happens the day it renders 47 thousand and my real answer is 4700? That number sits inside the plausibility envelope. My whole governance stack, budgets, alerts, chargeback, forecast, all assume the source of truth is at least directionally right.

I run a cost pipeline that produces the numbers a bunch of downstream controls consume. Once, on a chunked backfill against a partitioned billing dataset, I lost a partition-pruning predicate and the same pipeline scanned the same 146 gigabytes per chunk over and over. The bill told me before any dashboard did. It was the first time I trusted my bill more than my own instrumentation. Made me realize the instrumentation was never actually verified against the source of truth. Only against my expectations.

Every cloud treats these display faults as one-off gotchas per engine, per console, per service. No one publishes a trust model. The vendor page says the dashboard is your source of truth. The reconciled invoice says it is the source of truth. Nobody talks about how those two disagree in real time and which one your downstream automation is supposed to believe.

Genuine question for anyone running cost governance on a real bill. How do you build trust in the source of truth for the numbers your automation acts on?

- Do you reconcile dashboard against raw billing export daily, and if so what tolerances do you allow?

- Any programmatic sanity check for physically impossible values, or do you rely on humans noticing?

- If your budget alert or auto-shutdown fires on a display-side hallucination, what is your rollback path?

- Anyone actually run cost data through the same integrity discipline as production telemetry?


r/aws 12h ago

technical question What is happend to AWS - No details about gemma 4 in console...... but pricing page covers all

2 Upvotes

r/aws 6h ago

discussion Network engineer journey to Cloud

3 Upvotes

Cloud engineers, wanted to get your experience... I'm a network engineer with 15 years of experience with all kinds of on-prem network technologies, from NX-OS, load balancers, proxies, VMware, ACI. I'm currently working with NSX and AVI LB for a major bank. But with the Broadcom aquisition, VMware/NSX doesn't seem so appealing anymore, VMware jobs are very rare. I feel that I'm a niche that will die eventually and it's time to make a change. I have experience with Terraform and CI/CD pipelines, did some automation with Python vibe coding.

There are a lot of Cloud-related jobs and I like public cloud, I like to learn new stuff in general. I started to learn AWS and Azure. I got the SAA-C03 AWS Solution Architect Associate certification and now I'm learning to get the AZ-700 Azure Networking speciality. I applied to Cloud Network Engineer jobs but got rejected, probably due to missing on-the-job experience. At my current job I can't get any Public Cloud exposure. I did put in my CV a project in Github with Terraform standing up an AWS environment with ECS, load balancer, instances connecting over VPN to a VM in GCP.

How did you guys make it? It's the chicken and the egg... To get a job you need experience, but to get experience you need the job :)


r/aws 22h ago

general aws AWS cloud cohort

1 Upvotes

Does anyone know what this is? I got invited for it and i’m in college but idk if it’s worth it to do


r/aws 3h ago

general aws How do I get out of SES sandbox for Cognito (low volume)?

0 Upvotes

Hey all!

I built my auth stack around Cognito for my app, but I’m blocked from launching because I can't get out of the SES sandbox to send basic verification and password reset emails.

I submitted my sandbox exit request twice now, but I keep getting the same generic rejection email telling me they can't approve it to protect my deliverability, with zero actual feedback on what's wrong.

Before I give up on SES and hook up SendGrid or Resend, I wanted to see how others handle this:

  1. Is AWS auto-rejecting requests if the account is pretty new or low spend? (I’ve had an account for several years but very small scale, for personal projects)
  2. If you gave up on SES for Cognito, did you just use a Lambda trigger for another provider, or scrap Cognito entirely?

For context, here’s what I included in my requests:

  • Use case: Strictly transactional Cognito emails (verification, password resets, internal feedback notifications). No marketing emails
  • Recipients: Only users signing up directly on my app
  • Volume: Under 100 emails a day to start
  • Setup: Verified domain in us-east-1, Easy DKIM on, custom MAIL FROM subdomain, aligned SPF, Return-Path, and DMARC set up
  • Monitoring: Account suppression on, dedicated config set dumping bounce/complaint events to SNS, CloudWatch alarms configured (<0.1% complaint / <5% bounce).
  • Site: Terms and privacy policy are linked right on the login screen (with an opt-in box). The inbox is also monitored

Am I missing anything to get this approved? Appreciate any tips. Thanks!


r/aws 17h ago

discussion What’s it like being a Pro Serve Consultant?

0 Upvotes

I have an upcoming interview this week for a role.

Also, are all pro serve consultants mandated to be in the office 5 days a week (when not on the client site)?


r/aws 5h ago

technical question What are MSPs using for AWS EC2 backups in 2026?

7 Upvotes

Hi everyone, We are currently managing a few AWS EC2 environments that were previously backed up using Acronis (via eFolder) as part of a standardized setup. with that integration no longer being viable for us, we're looking at replacement options.

For on prem environments we typically use solutions like Datto and Replibit, but they dont translate well to cloud native workloads.

I'm trying to understand what MSPs are actually using today for EC2 backups. are most people relying on AWS native tooling( EBS snapshots, AWS backup, S3 based policies) or are third party platforms still the preferred route?

We also have a couple of instances running MS SQL so proper application aware backups or database consistent snapshots are important.

I've considered building a more AWS native setup but id prefer something that doesnt require a heavy custom scripting or ongoing CLI based automation management.

Would appreciate any real world setups or recommendations that are working well in production


r/aws 3h ago

billing Is my understanding correct, that AWS SES is more expensive now?

Thumbnail gallery
7 Upvotes

AWS SES introduced Pricing Plans:
https://aws.amazon.com/ses/pricing/

The first screenshot is inside my dashboard, which shows I'm in the "a la carte" pricing.

Why do they say "go to a pricing plan to have discounts",
when to me it seems that sending emails is going to be more expensive?????

I only send emails with SES, nothing else I do with SES.
It seems to cost me $0.10/1000 right now. Why would I go to a plan and pay more per 1000 emails ?!

And the worst thing is that it seems that once I go to a pricing plan, I cannot go back to "a la carte" (otherwise I'd see 4 boxes instead of 3, I assume).

Am I seeing something wrong here?


r/aws 6h ago

technical resource Two surprise AWS bills later, I built the ecommerce project I wish existed

Thumbnail gallery
0 Upvotes

I've been burned twice by AWS tutorials that glossed over billing.

The first time, I followed a Bedrock RAG walkthrough step by step. Everything worked, felt great.

Then I opened Cost Explorer two days later.

$70.

OpenSearch Service had been running the whole time and it was never mentioned in the tutorial. Not the cost, not the hourly rate, nothing.

The second time was a SageMaker project.

I built it, tested it, and deleted everything I could see.

Thought I was clean.

I wasn't.

Some underlying resources SageMaker creates don't get deleted when you tear things down, and they aren't always obvious to find.

I eventually tracked them down using the AWS FinOps Agent (still in preview). By then I had an $11 bill sitting there.

Neither of these happened because AWS is expensive.

They happened because the cost side of cloud tutorials is usually an afterthought.

And if your free tier has expired, that gap in knowledge makes you hesitate to build anything because you genuinely don't know if following a tutorial will cost you $2 or $70.

So I decided to build the project I couldn't find.

What it is

NordicCart is a fictional small ecommerce retailer that starts with a very common setup:

One EC2 instance.

No redundancy.
No CDN.
No scaling.
No cost visibility.

Over 7 days, the project modernizes it using AWS services that a real small retailer would actually need.

What gets built

  • Custom VPC across two Availability Zones
  • Application Load Balancer + Auto Scaling Group (actually tested with a live load test)
  • DynamoDB as the data layer instead of RDS (with a design doc explaining why)
  • Lambda + API Gateway + Amazon Bedrock AI support assistant with conversation memory
  • Static frontend hosted on S3 and delivered through CloudFront
  • AWS Budgets + SNS email alerts before spending exceeds $5
  • CloudFormation stack that deploys everything in around 20 minutes

Bonus:

A working storefront where you can add items to cart, check out, receive an order ID stored in DynamoDB, then ask the AI assistant about your order and it can actually retrieve the information.

Every resource is tagged from day one:

project, environment, component, owner.

The final day includes a tag audit specifically because of the SageMaker cleanup issue I ran into.

Cost breakdown

Bedrock is the only real usage-based cost.

There is no free tier for model invocations, but with short prompts and light testing the entire build stays under $1.

The project was built on a free tier account and stayed within that range.

The ALB is free for the first 12 months of a new AWS account. After that, build it, test it, capture your screenshots, then tear it down. A few hours of testing costs cents.

Everything else is either Always Free or covered by the 12-month free tier at this project's scale.

I also deliberately skipped the NAT Gateway.

That one decision saves around $32/month, and I redesigned the VPC architecture around that constraint.

Repo

The goal was not just to build something that works, but to document the decisions that affect your AWS bill.

There are:

  • CLI and console paths for every step
  • explanations for when I avoided the "textbook" architecture because of cost
  • trade-off discussions
  • troubleshooting notes from mistakes I actually hit

Repo:
https://github.com/BTAG16/nordiccart-aws

Would love feedback, especially from people who have built similar AWS portfolio projects.

What AWS service has surprised you with a bill before?


r/aws 15h ago

article Compiling and running a pre-trained LLM on AWS Inferentia accelerator

Thumbnail pooria.co
5 Upvotes

In this tutorial, we are going to compile and run a small llama architecture model on an EC2 instance and if we manage to pass the compilation and inference test, it means our model is compatible.

Source code in Github: https://github.com/p0o/run-models-in-aws-inf2-ml-accelerator