Help AWS customers get fast, confident answers when their network traffic misbehaves.
Hyperplane is the distributed network virtualization platform behind AWS services such as NAT Gateway, Network Load Balancer, Gateway Load Balancer, AWS PrivateLink and Elastic File System. It processes trillions of packets a day for millions of customer workloads. At that scale, finding the root cause of a customer issue is a hard technical problem. A single flow can cross several services, each with its own telemetry, and the answer is often hidden in a few packets among billions.
We're investing in making those answers faster and more reliable. You'll build debug tools that both on-call engineers and AI agents can use. They'll correlate evidence across services for a single flow, recognize known failure patterns at once, and define the telemetry our dataplane teams build next.
This is a high-ownership role with a lot of room to shape what we build and how. You'll work real customer escalations alongside building the tools, because the best debug tools come from people who have done the debugging. We expect you to use generative AI as a core part of how you work: to ramp up quickly in large, unfamiliar C, C++ and Rust codebases, to change existing systems safely, and as a building block in the diagnosis tools you create.
Ready to make a durable mark on the computer industry? Come join us for this unique opportunity that blends leadership, technology, and accelerating customer growth!
Key job responsibilities
• Investigate complex customer network issues across Hyperplane and partner services (load balancers, firewalls, endpoints), from packet captures down to dataplane source code, and turn each investigation into a reusable tool or pattern
• Design and build tools that pull together evidence for a customer flow across multiple services and return a structured verdict quickly
• Build and evaluate AI-agent-based diagnosis capabilities (for example on Amazon Bedrock AgentCore), including measuring whether the agents actually get it right
• Build and maintain a library of known failure patterns drawn from past escalations, and make it something both people and AI agents can use
• Define requirements for new telemetry (drop attribution, flow-level visibility), back them with data from real tickets, and influence dataplane teams to deliver them
• Use AI-assisted development to understand and change existing production code across several services and languages, while keeping a high bar on correctness and operational safety
• Publish and document the drop reasons, telemetry and debugging practice so knowledge stops living in individual engineers' heads
• Mentor engineers across the Hyperplane org on debugging technique, and share what you learn through docs, talks and tooling
A day in the life
You might start the morning on a customer escalation where traffic is being silently dropped. You use an AI assistant to trace a code path you've never seen in a dataplane component, check the result against flow data, and reach a verdict. In the afternoon you turn what you learned into a new failure pattern and a check the next on-call engineer gets for free, then review telemetry requirements with a Seattle dataplane team. Some weeks are mostly building and some are mostly digging in. Both are the job.
About the team
AWS Infrastructure Services designs, plans, delivers, and operates AWS's global infrastructure. We keep the cloud running. We support every AWS data center and the servers, storage, networking, power, and cooling that keep customers connected to the innovation they rely on. We tackle hard problems with thousands of supply chain variables and need talented people to help.
You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, and operations managers. Together we deliver strong safety, security, and seemingly infinite capacity at the lowest cost, in an inclusive culture where you own bold ideas.
Basic Qualifications:
- 5+ years of non-internship professional software development experience
- 5+ years of programming with at least one software programming language experience
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience as a mentor, tech lead or leading an engineering team
Preferred Qualifications:
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit
https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. As a total compensation company, Amazon's package may include other elements such as sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life & AD&D insurance), Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and other resources to improve health and well-being. We thank all applicants for their interest, however only those interviewed will be advised as to hiring status.
CAN, BC, Vancouver - 150,700.00 - 251,700.00 CAD annually