Jobs in Sussex County, NJ

30verified openings
Filter by
No filters selected
planetnetworks Verified 2d ago

Senior VoIP Operations & Reliability Engineer (Carrier-Class Voice Platform)

Newton, New Jersey, 07860-1766, United States On-site

Salary not listed
PaySalary not listed
TypeFull-time
Work settingOn-site
Verified listing

JobFig found this opening at its original source and checks that it remains available.

About the role

Our software team is building a next-generation carrier-class voice platform. They are strong programmers, but they are not experienced operators, and there is a world of difference between code that works and infrastructure that stays up under real carrier load. We need a seasoned operator to close that gap and work hand in hand with the development team. You are the person who has actually run this kind of system in production. You know the failure modes that do not show up in a code review, the things that break at 2am, and what it really takes to keep customers from ever noticing. Your job is to bring that operational reality into the platform from the inside: pairing with the programmers as they build, making sure the design can be operated, and then owning the platform in production with zero customer downtime. There is an architectural side to this. You will sit in design reviews and push the team toward decisions that are operable, resilient, and testable, not just elegant in code. But the core of the role is operational: you are the experienced hand who keeps every system up, who owns every failure scenario end to end, and who instills operational discipline in a team by being on-call and training  juniors to handle any incidents. You should be equally comfortable pairing with a developer to make a service observable and failure-aware, and at 3am driving an incident to resolution. We need that judgment, with years of real VoIP operations behind it. In the meantime, this is not a future-only role. We already run a live Kamailio and Asterisk production system carrying real customer traffic today, and your first and most immediate mandate is to help harden it: shore up its reliability, close its failure gaps, and keep it solid while the next-generation platform is being built. Day-to-day production stability of the current system comes first.

What you'll bring

  • Years of senior, hands-on experience operating and reliability-engineering production VoIP systems at carrier scale.
  • Deep, protocol-level command of SIP: dialogs, transactions, registration, NAT scenarios, SDP negotiation, forking, and the failure modes that surface only under load.
  • Expert-level Kamailio and/or OpenSIPS: routing logic, dispatcher and load balancing, registrar and usrloc, dialog and topology modules.
  • Expert-level Asterisk: PJSIP stack, dialplan, ARI/AMI, bridging and media handling, and its role as an application and media server behind a SIP proxy.
  • Media plane fluency: RTP, SRTP, RTSP, RTCP, codecs (G.711, G.729, Opus), transcoding, jitter, and the link between QoS marking (DSCP) and call quality.
  • A demonstrated track record of designing for and operating reliability, scalability, and fault tolerance in carrier-class environments (five-nines thinking, failure-domain isolation, blast-radius control).
  • Hands-on reliability engineering practice: SLOs and error budgets, incident command, postmortems, runbooks, and DR testing.
  • Performance and failure testing tooling: sipp for load and call modeling, fault injection and chaos tooling, and SIP troubleshooting with sngrep and Wireshark.
  • Observability depth with Homer/HEP, plus metrics and alerting stacks (for example Prometheus, Grafana, or equivalent).
  • Strong Linux operations and automation skills (Python, Lua, shell), and comfort with infrastructure-as-code and CI/CD pipelines.
  • RADIUS/Diameter integration for AAA, and experience with provisioning and subscriber management.
  • Fraud and security operations: detecting and stopping toll fraud, SIP scanning, and registration attacks.