Infrastructure Engineer
We are seeking an infrastructure engineer who owns a real on-premises estate and is comfortable stepping into cloud and automation work when the team needs it.
This is not a cloud-native DevOps role. The core of the job is networking, firewalls, virtualization, and storage across multiple datacenters. Cloud, Terraform, and CI/CD sit alongside that as a valuable second skill set rather than the main event. You will work day-to-day alongside the product infrastructure manager, take real ownership of projects, and act as the technical escalation point for infrastructure incidents.
About First Factory
We are a software development company with over two decades of experience, boasting a dynamic team of 175+ professionals actively engaged in diverse projects across various industries. We invite you to join us on this journey as we thrive and embrace fresh challenges.
Key Responsibilities
Configure, maintain, and troubleshoot the network estate: Cisco switches and routers (Catalyst series), firewalls including HA clusters, SD-WAN, IPsec VPN tunnels, datacenter switching (Mellanox experience a plus), and diagnose ISP circuit-level issues to manage carrier fault tickets.
Harden and secure the infrastructure: lock down switch and device management, control VPN access, and keep external actors out of the estate.
Administer VMware vCenter and ESXi environments across multiple sites using PowerShell and PowerCLI for automation.
Administer storage arrays (provisioning, performance monitoring, troubleshooting; Pure Storage, Dell SAN, or Storage Spaces Direct) and own backup and disaster recovery workflows, including restores using Nakivo, Wasabi, and Rclone.
Manage server hardware lifecycles remotely: firmware, hardware diagnostics, vendor/lease relationships, and directing colocation remote hands for replacements, power cycling, and console access.
Own SSL certificate renewal and deployment back onto platforms and network devices.
Implement and maintain monitoring coverage, alerting, and dashboards for legacy and cloud infrastructure using LogicMonitor and Datadog.
Provision and maintain AWS infrastructure (EC2, ECS Fargate, RDS, VPC, IAM, SNS, Systems Manager, Secrets Manager, WAF, Route 53, S3) and keep connectivity between on-premises data centers and the cloud secure and reliable.
Write, review, and maintain Infrastructure as Code (IaC) using Terraform for AWS, VMware, and network configurations; develop CI/CD pipelines in GitHub Actions (YAML, OIDC).
Serve as a technical escalation point for infrastructure incidents, driving full-stack troubleshooting across network, storage, virtualization, cloud, and application layers.
Join the on-call rotation, take ownership of alerts end-to-end, and advise the development team on infrastructure/platform architecture.
Requirements
Proficiency with Cisco switching and routing (Catalyst series), VLAN configuration, and routing protocols.
Hands-on firewall configuration experience (FortiGate preferred; Cisco ASA, Palo Alto, or comparable platforms welcome).
Practical experience administering storage arrays, RAID levels, and disk layouts (Pure Storage, Dell SAN, or Storage Spaces Direct).
Solid background managing VMware environments (vCenter, ESXi).
Experience owning backup and DR: retention, rotations, offsite copies, and restores you have actually performed.
Experience managing physical server estates, firmware/hardware diagnostics, and directing third-party remote hands.
Strong AWS administration skills (EC2, ECS Fargate, RDS, VPC, IAM, SNS, Systems Manager, Secrets Manager, WAF, Route 53, S3) with a working knowledge of Azure.
Ability to configure end-to-end monitoring, alerting, and custom dashboards (LogicMonitor, Datadog, or comparable).
Judgment to read scripts, understand blast radius, and validate before running against production.
Highly self-directed with a proven ability to take full ownership of complex technical problems through to resolution without close supervision.
Sharp troubleshooting instincts across network, virtualization, storage, cloud, and application layers.
Clear written and verbal English communication skills for technical documentation, status updates, and vendor management.
Nice to have
Terraform and Infrastructure as Code (AWS, VMware estate, firewall/switch configs in GitHub).
CI/CD pipeline work in GitHub Actions: deployment/workflow YAML and OIDC connections between GitHub and AWS.
Containerization and orchestration (ECS/Fargate; Kubernetes experience translates directly).
PowerShell, PowerCLI, and Bash scripting for VMware and operational automation.
Mellanox datacenter switching experience.
Specific backup tooling experience: Nakivo, Wasabi (S3-compatible cloud storage), and Rclone.
Basic Postgres administration (database/user creation, backup/restore, identifying long-running queries).
Production on-call and incident management experience.
Jira and Confluence for ticketing and documentation.
Experience with managed hosting platforms and migrating workloads off them into AWS.
- Department
- Software Engineering
- Role
- Infrastructure & DevOps Engineer
- Locations
- Heredia
- Remote status
- Hybrid
About First Factory
For over 25 years, First Factory has been a place where collaborative excellence meets modern technologies. We’re a strong team building exceptional software solutions from Costa Rica and LATAM for primarily US-based clients. With industry-low turnover, top eNPS globally, and 5 consecutive Inc. 5000 awards, we foster an environment where talented engineers thrive on challenging projects using modern tech stacks.