Why Castelion, Why Now
Castelion is moving incredibly fast to develop and deliver advanced defense systems at a time when execution matters more than ever. We believe focus, ownership, and excellence are decisive advantages - and we're building a world-class team to turn bold ideas into real capability.
This is a rare opportunity to join at an early stage, where you'll have significant ownership, collaborate with exceptional teammates, and make a direct, measurable impact on our mission and the future of the company - regardless of your function.
Site Reliability Engineer
We are seeking a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.
This role is the missing reliability piece of an existing high-performing engineering organization. You will work across DevOps, Cloud, Software, Security, Test, and IT to identify reliability risks, diagnose failures that cross system boundaries, and drive corrective actions to resolution. You will be expected to understand and improve existing systems rather than defaulting to replacement, using new technology when it solves a demonstrated reliability, scalability, or operational problem.
Responsibilities
- Establish meaningful reliability, availability, latency, capacity, and recovery expectations for critical engineering services, with measurable health indicators and useful alerts.
- Lead deep technical investigations and incident response across application, Linux, networking, storage, Kubernetes, cloud, and other system boundaries; collect evidence, separate symptoms from root causes, and drive incidents through resolution.
- Build and improve monitoring and diagnostic systems that detect problems before users report them and provide engineers with the information needed to quickly understand and resolve failures.
- Analyze system performance and capacity across compute, memory, storage, networking, connections, and other constrained resources; identify operating limits and address issues through the simplest effective solution, whether optimization, additional capacity, scaling, caching, configuration changes, or architectural improvements.
- Drive evidence-backed root cause analysis and postmortem actions for significant incidents, ensuring corrective and preventive actions are implemented and verified to reduce recurring failures.
- Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve reliability problems that cross team boundaries, providing technical leadership without attempting to own every component involved.
- Understand, operate, and incrementally improve systems built by other engineers, balancing reliability and operational value against existing architecture, constraints, and engineering practices.
- Participate in the on-call rotation for critical engineering services, providing first-response triage, escalation, and follow-up for recurring reliability issues.
Basic Qualifications
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
- Demonstrated experience debugging complex production failures across multiple system layers and driving investigations from the first symptom to an evidence-backed root cause and lasting corrective action.
- Strong Linux systems expertise, including CPU, memory, storage, networking, processes, sockets, and system services, with a strong understanding of performance and capacity concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
- Experience building and operating observability, monitoring, alerting, and incident response systems, with the ability to distinguish between mitigation, workaround, corrective action, and preventive action.
- Strong networking and application fundamentals, including TCP, TLS, HTTP, DNS, reverse proxies, load balancers, connection states, and timeouts; able to investigate application runtime behavior such as threads, connection pools, file descriptors, memory, or garbage collection when the evidence points there.
- Demonstrated ability to work effectively within existing systems and across engineering organizations, asking why a system was designed a certain way and improving it based on measurable reliability and operational needs rather than defaulting to rewrites or replacement.
Castelion offers a generous benefits package. Please refer to the bottom of our Careers page for more details.
Other Duties
Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.
Additional Eligibility Requirements
This position may require access to classified information or restricted U.S. Government sites, systems, or information, as determined by the Company and/or applicable U.S. Government requirements. If the position is so designated, your employment in the role may be contingent upon your ability to obtain and maintain the required U.S. Government security clearance or other government authorization, and to satisfy any citizenship or other eligibility requirements imposed by applicable law, regulation, executive order, or government contract requirements. You will be notified if and when such requirements apply.
Affirmative Action/EEO Statement
Castelion is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.
Castelion is committed to providing reasonable accommodations to qualified individuals with disabilities throughout the application and hiring process. If you require a reasonable accommodation to complete an application, participate in the interview process, or otherwise participate in the hiring process, please contact [email protected]. Requests for accommodation will be considered on an individual basis and handled in accordance with applicable law.
Castelion is committed to fostering a workplace where employment decisions are based on qualifications, business needs, and the ability to perform the essential functions of the role, with or without reasonable accommodation.
EAR/ITAR Requirements
This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.