Full Time
Remote
Site Reliability Engineer

Overview
Trajeco Industries builds the custom software that runs manufacturing and aerospace operations, and we operate that software ourselves in production. As a Site Reliability Engineer, you’ll be at the center of that work, keeping our software running when a customer’s shift depends on it.
The role
You’ll monitor, troubleshoot, and improve the reliability of our production systems from day one, working across incident response, observability, and capacity planning. This is hands-on engineering: you’ll be paged for real incidents as often as you’re improving systems to prevent the next one.
What you’ll do
Build and maintain monitoring and alerting for our production software, lead incident response, and work with our engineering teams to fix the root cause behind every recurring issue. You’ll own uptime from first alert through resolution and iterate as our systems scale.
What you’ll bring
Several years of hands-on experience operating production systems at scale, plus a strong grounding in observability and incident response. You’re comfortable owning reliability end to end and know the difference between a fix that works once and one that holds up long-term.
Bonus points
Experience keeping manufacturing or industrial software online, chaos engineering, or building systems with strict uptime requirements. Familiarity with capacity planning and how it feeds long-term stability is a real plus.
Why Trajeco Industries
We build our own reliability practice because uptime is the product as much as the software itself. You’ll work on real systems that run on real production floors, alongside a team that designs, builds, and operates everything in-house. We build software in the U.A.E., and we build it to last.
How to apply for this role
If you think you're the right for us, we'd love to hear from you! Please click the link below to apply for this role.