Remote Operations and Support Lead
- Published on 10/03/2026
- Hartford (CT003)
- To be defined
Description:
Hydra Host operates mission-critical AI infrastructure where customer success depends on operational excellence. This role combines Customer Support Leadership with Infrastructure Operations Management , serving as the operational hub between customers, engineering, deployment teams, hardware vendors, and AI Factory partners. You'll own customer-facing operational support while building the internal processes that keep our NeoCloud platform running efficiently. Whether responding to a GPU outage, coordinating infrastructure deployments, managing vendor escalations, improving SLAs, or building scalable operational workflows, you'll ensure both our customers and our infrastructure perform at the highest level. This is not a traditional support management position. This is an operations leadership role responsible for the daily execution, reliability, and continuous improvement of Hydra Host's AI Factory platform. What You'll Do Lead Customer Support Operations Develop and lead Hydra Host's customer support organization supporting: Design and build Hydra Host’s customer support organization from the ground up Enterprise AI customers Interface with the Machine Learning engineering teams GPU infrastructure customers AI Factory operators Data center partners Build a high-performing support organization that delivers exceptional customer experiences while maintaining enterprise-grade service levels. Own Operational Excellence Drive the day-to-day operational health of Hydra Host's NeoCloud platform by coordinating activities across engineering, infrastructure, vendors, and customer-facing teams. Ensure infrastructure operates reliably while continuously improving operational efficiency. Lead Major Incident Management Serve as the Incident Commander during production-impacting events. Coordinate engineering, networking, infrastructure, vendors, and customers to resolve: GPU cluster failures Network outages Hardware failures Firmware issues Storage performance degradation Infrastructure capacity constraints Customer-impacting production incidents Own customer communications throughout incident response while driving rapid resolution and post-incident improvements. Build Scalable Support Operations Design and implement: Ticketing systems Escalation procedures Knowledge management Support automation AI-powered support tools On-call rotations Operational playbooks Customer communication standards Create a support organization capable of scaling alongside Hydra Host's rapid growth. Drive AI Factory Operations Partner with deployment, engineering, and data center teams to coordinate: Infrastructure deployments Rack turn-up GPU cluster readiness Network activation Capacity planning Maintenance scheduling Production acceptance Operational readiness reviews Help ensure AI Factory infrastructure is deployed efficiently and operates reliably. Vendor