About the Role
About the Role
As a Systems Development Engineer in Operations Infrastructure Services at Amazon, you will design and deliver infrastructure monitoring applications that enhance global operations. This role involves building scalable monitoring solutions, applying AI technologies, and collaborating with various teams to ensure system reliability and operational excellence.
What You'll Do
- Design, build, and operate scalable monitoring infrastructure on AWS that ingests and processes high-volume device telemetry and network topology data.
- Own systems end-to-end across monitoring tooling and data pipelines, from build through deployment, operation, and ongoing maintenance.
- Partner with infrastructure and operations stakeholders to grasp monitoring gaps and translate them into reliable, automated solutions that reduce manual intervention.
- Apply AI/ML and generative AI techniques to improve detection quality, reduce alarm noise, and streamline operational workflows.
- Raise the bar on system reliability, operational excellence, and automation through thoughtful design, thorough testing, and continuous improvement.
A Day in the Life
You might start by reviewing a design for integrating metrics from a new device type being onboarded across all Robotics buildings worldwide, then inspect runtime metrics to tune metric collection and quality before shipping a service improvement our operators experience immediately. Our team values partnership, quick feedback loops, and clear ownership.
Benefits
- Medical, Dental, and Vision Coverage.
- Maternity and Parental Leave Options.
- Paid Time Off (PTO).
- 401(k) Plan.
About The Team
We're a close-knit, agile team that owns infrastructure monitoring for OIS within Amazon Robotics and ships to a worldwide operational fleet. We care deeply about the craft of software and foster each other's growth. Our vision centers on delivering reliable, intelligent monitoring that helps operators grasp their infrastructure at scale.
Requirements
AWS experience
Experience in designing and operating scalable solutions on AWS.
AI/ML knowledge
Familiarity with AI and machine learning techniques to enhance monitoring solutions.
Telemetry processing
Ability to process high-volume device telemetry and network topology data.
Collaboration skills
Strong partnership skills to work with infrastructure and operations stakeholders.
Nice to Have
Experience with generative AI techniques is a plus.
Familiarity with monitoring tools and data pipelines.
Benefits
Health coverage
Comprehensive medical, dental, and vision coverage.
Parental leave
Options for maternity and parental leave.
Paid time off
Generous paid time off policy.
Retirement plan
401(k) plan to help you save for retirement.