We are looking for a highly skilled Senior NOC Engineer with 6 to 8 years of experience to join our 24/7 Network Operations Center. In this role, you will be the second-line anchor for a high-concurrency production VoIP PBX platform.
You will be responsible for real-time proactive monitoring, incident triage, and runbook remediation of our real-time communication (RTC) stack. This includes managing Kamailio edge proxies, FreeSWITCH media engines, and PostgreSQL databases, all running on Google Cloud Platform (GCP) Compute Engine VMs with dedicated static IP carrier trunks. You will manage P1/P2 alerting flows, coordinate tier-3 engineering escalations, and protect core telephony uptime.
🔧 Key Responsibilities
- 24/7 Eyes-on-Glass Monitoring: Proactively monitor platform health dashboards across Grafana, GCP Cloud Monitoring, and Homer (SIPcapture) to detect signaling drops, media quality degradation, or infrastructure bottlenecks.
- Incident Triage & Response: Act as the Tier 2 technical first responder for platform alerts. Accurately classify incidents vs. outages according to business logic (P1 to P5 matrix).
- Runbook Execution & Remediation: Execute specialized technical runbooks to resolve complex VoIP faults (e.g., clearing stuck SIP registrations, flushing PgBouncer connection pools, safely draining/bypassing misbehaving FreeSWITCH instances).
- VoIP Media Quality Analysis: Troubleshoot voice quality spikes (Jitter >50ms, Packet Loss >2%) and analyze Mean Opinion Score (MOS) degradation using Homer SIP ladder diagrams.
- Carrier Interconnect Monitoring: Manage links tied to static public IP carrier trunks. Monitor SIP OPTIONS pings and immediately troubleshoot authentication issues (SIP 401/403) or error loops (SIP 503).
- Escalation & Collaboration: Gather raw packet traces (sngrep / Wireshark) and technical context during major incidents to smoothly escalate unresolved issues to Tier 3 DevOps, VoIP Core, or Database Engineers.
- Leadership & Mentorship: Act as the Shift Lead during critical hours, mentoring junior NOC operators, managing incident communication channels, and refining technical runbooks.
🎯 Technical Requirements & Qualifications
- Experience: 6 to 8 years in a 24x7 technical operations center environment, with a heavy emphasis on VoIP, SIP Signaling, and Cloud Infrastructure.
- VoIP Stack Expertise: Direct, hands-on experience troubleshooting and managing Kamailio (or OpenSIPS) and FreeSWITCH (or Asterisk) configurations and logs.
- Protocol Proficiency: Deep understanding of SIP signaling, RFC standards, SDP negotiation, and RTP/RTCP media streaming behaviors.
- Cloud Platform Skills: Strong proficiency with Google Cloud Platform (GCP) core services, specifically Compute Engine Virtual Machines (VMs), VPC Networking, Cloud Firewall, Layer 4 Network Load Balancers, and Static External IP architectures.
- Database Foundations: Competence working with PostgreSQL (specifically Cloud SQL environments); ability to analyze active connection volumes, understand PgBouncer states, and identify query blockages during high call volume.
- Observability Tools: Hands-on experience navigating Homer (SIPcapture), Wireshark, sngrep, Prometheus, and Grafana alert configurations.
- Tooling: Familiarity with Incident Management systems such as PagerDuty, Opsgenie, Jira Service Desk, and Slack integrations.
🌟 Soft Skills
- Calm Under Pressure: Proven ability to handle live, high-stress P1 Critical Outage scenarios with structure and focus.
- Clear Communicator: Strong verbal and written English communication skills to act as a bridge between technical engineering pods and customer-facing support teams.
- Problem Solver: Analytical mindset capable of isolating a network hop failure vs. a software application core dump.