Job Overview
Senior DevOps Engineer
Vexere is a technology company aiming to revolutionize the travel and transportation industry in Vietnam. We are pioneering a comprehensive digital transformation of the transport industry with diverse web applications and latest technology apps. We are available on all super apps such as Grab, Momo, and Zalopay. Covering diverse vehicles from low-priced to luxury buses, with 1,000 bus operators on all over 3,000 routes, Vexere is the largest online inter-city bus booking platform that empowers millions of travelers to make their life journeys happier.
Want to secure a large-scale cloud-native platform with 100+ production services?
Want to turn DevOps strategy into practical controls across Kubernetes, CI/CD, cloud infrastructure, secrets management, observability, and public-facing systems?
Come to us to build and improve security for a modern microservices ecosystem using GitLab CI, Flux CD, Kubernetes, Vault, Cloudflare, OpenTelemetry, SigNoz, ELK, VictoriaMetrics, and other production-grade technologies. This is a hands-on engineering role, not a pure compliance or advisory role. You will work directly with a small DevOps team that owns both platform operations and security improvements, with autonomy to build scripts, CI templates, dashboards, runbooks, policies-as-code, and practical controls that reduce real production risks.
Working model: remote-first or hybrid, with planned offline meetings twice each week.
Reporting line: DevOps lead
Environment
- 3 cloud providers: Google Cloud, Viettel Cloud, and CMC Cloud.
- Around 50 million requests per day.
- 100+ production services on Kubernetes.
- Multiple WordPress deployments on Kubernetes.
- Production databases across MSSQL, MongoDB, PostgreSQL, ClickHouse, and MySQL.
- GitLab CI, Flux CD, Terraform/Terragrunt, Ansible, Vault, OpenTelemetry, SigNoz, ELK, VictoriaMetrics, and Grafana.
- Self-hosted stateful infrastructure and high-impact production workloads.
Responsibilities
Platform Reliability And Incident Response
- Operate Kubernetes clusters, ingress, DNS, networking, storage, workload scheduling, and shared runtime components.
- Investigate incidents using metrics, logs, traces, Kubernetes events, infrastructure state, and application signals.
- Improve alerts, dashboards, escalation paths, and runbooks for critical systems.
- Participate in production escalation, including after-hours incidents when required.
- Convert repeated incidents into automation, standards, capacity changes, or documented prevention work.
Cloud Infrastructure And GitOps
- Build and maintain infrastructure using Terraform/Terragrunt, GitLab CI, Flux CD, and Ansible.
- Standardize cloud patterns across Google Cloud, Viettel Cloud, and CMC Cloud where practical.
- Manage network connectivity, firewall rules, VPN, DNS, service exposure, certificates, and cloud resource lifecycle.
- Improve rollback, break-glass access, production change review, and GitOps recovery procedures.
- Review infrastructure changes for reliability, security, cost, and operational impact.
CI/CD And Developer Enablement
- Maintain reusable GitLab CI templates and deployment patterns.
- Help application teams adopt standard patterns for configuration, secrets, rollouts, observability, and production readiness.
- Reduce one-off support through self-service workflows, checks, and concise documentation.
Database, Data, And Stateful Operations
- Support operational work around MSSQL, PostgreSQL, MongoDB, MySQL, ClickHouse, Redis, Kafka, and Elasticsearch.
- Improve replication, capacity planning, performance investigation, and recurring data-pipeline operations.
- Coordinate high-risk data changes with service owners and database stakeholders.
- Maintain runbooks for lag investigation, restore testing, failover, and data incident response.
High Availability And Disaster Recovery
- Improve availability patterns for workloads, ingress, databases, caches, queues, storage, and cloud infrastructure.
- Validate backup, restore, failover, rollback, and recovery procedures for critical systems.
- Run disaster recovery drills and document RTO, RPO, and remediation gaps.
- Review architecture changes for single points of failure, capacity risk, zone dependency, and recovery impact.
Automation And AI Workflows
- Build scripts or tools for inventory, drift detection, alert review, cost review, backup checks, and recurring operations.
- Use AI-assisted workflows for log analysis, incident summaries, runbook drafts, infrastructure review, and operational reporting.
- Keep human approval for production changes, database recovery, security exceptions, and disaster recovery execution.
Required Experience
- Production Kubernetes experience covering deployments, services, ingress, autoscaling, resource limits, troubleshooting, and recovery.
- Hands-on Linux, networking, DNS, TLS, load balancing, firewall, VPN, and production debugging experience.
- Cloud infrastructure experience. Google Cloud is preferred; multi-cloud experience is a strong plus.
- Infrastructure as code experience with Terraform or Terragrunt.
- CI/CD experience. GitLab CI is preferred.
- GitOps or Kubernetes deployment automation experience. Flux CD or Argo CD is preferred.
- Observability experience with logs, metrics, traces, alerts, dashboards, and incident analysis.
- Scripting or automation ability using Bash, Python, Go, or similar tools.
- Production experience with at least one database, cache, queue, or search system.
- Ability to write clear runbooks, incident notes, technical decisions, and async updates.
Operating Style
- Senior enough to own problems end to end without waiting for detailed task breakdown.
- Comfortable working in a small DevOps team where platform, cloud, security, database, observability, and support work compete for capacity.
- Prioritizes production safety, automation, documentation, and repeatable standards.
- Communicates clearly in writing and works effectively with distributed engineering teams.
- Can explain tradeoffs to developers, product teams, and leadership without hiding behind tool names.
- Accepts operational responsibility during incidents.
Preferred Signals
Certifications are optional. Equivalent hands-on production delivery is more important.
- Has operated Kubernetes for a production microservices platform with meaningful traffic.
- Has reduced alert noise, recurring incidents, manual changes, or deployment failures.
- Has built reusable CI/CD templates, GitOps workflows, Terraform modules, runbooks, or internal tools.
- Has handled incidents involving Kubernetes, cloud networking, databases, Kafka, Redis, Elasticsearch, or observability systems.
- Has performed backup and restore testing, disaster recovery planning, failover drills, or postmortems.
- Has worked at a startup, marketplace, fintech, ecommerce, travel, logistics, SaaS, or other high-traffic online platform.
- Has practical security experience around IAM, Vault, Kubernetes hardening, CI/CD controls, secrets rotation, or public endpoint protection.
- Has used AI tools responsibly for investigation, documentation, code review, automation, or operational analysis.
BENEFITS
Competitive Compensation & Bonuses
- Attractive salary package with performance-based monthly bonuses.
- 100% salary during the probation period.
- 13th-month salary to ensure financial stability.
- Performance reviews twice a year with opportunities for salary adjustments.
Flexible Work & Growth Opportunities
- Hybrid working model with remote work and offline in-person meetings twice each week.
- Work directly with the DevOps Lead and collaborate with experienced CTOs, architects, and tech leaders at Vexere.
- Codex AI access to support research, analysis, automation, and documentation work.
- Opportunity to secure a cloud-native microservices platform with meaningful production scale.
- Startup environment with low bureaucracy, no micromanagement, and direct ownership.
- Freedom to research, test, and apply new technologies when they provide practical value.
- Build your public profile through publishing technical articles, contributing to open-source projects managed by Vexere, and joining tech talks or industry events.
- Attractive stock options for dedicated developers and team leads.
Exclusive Employee Perks
- Up to 30% discount on bus tickets for Vexere employees and their families.
- 12 days of annual leave, with 1 extra day off every 3 years, convertible into salary.
- Vibrant and dynamic work environment with a friendly, supportive team.
- Training sessions on negotiation, communication, work management, interpersonal skills, and software technology.
- Exciting company activities: annual trips, team-building events, year-end parties, and more.
What factors can help you trust Vexere?
Currently, Vexere is the largest online bus ticketing platform in Vietnam with more than 1000 inter-city bus companies, covering over 3,000 domestic and cross-border routes to help users find bus information and buy tickets online easily. At the same time, Vexere also provides reviews of passengers who have traveled these buses. Passengers can simply select their favorite seats, pay online, or pay in cash at convenience stores across the country.
Vexere Bus Management System (BMS) is used by more than 700 bus companies all over the country, helping bus companies to modernize management with effective management, optimized revenue, and costs. This is a revolution in the inter-city bus industry.
Over 5,000 agents are using Vexere Agent Management System to increase ticket sales and serve their customers better.
Vexere’s efforts have been recognized and accompanied by famous institutional investors such as Woowa Brothers, CyberAgent Ventures, BonAngels, Pix Vine Capital, Spiral Ventures, Access Ventures, and NCore Ventures. Two investors of Vexere are Decacorns with valuation above USD10B.
For more information, please contact us via:
- Send your Resume to email: careers@vexere.com, with title: Fullname – Applied Position
- Phone/ Zalo: 0966 197 741 ( Mr. Anh)
- Office Location: Vexere Trading Service Co., Ltd – 2nd floor – Building H3, 384 Hoang Dieu, Ward 6, District 4, HCMC
- Working Hours: Hybrid Working 8:30 am – 6:00 pm from Monday to Friday, and Saturday morning.
