Company Description
Engineering the AI-powered enterprise. With AI and cloud-native solutions, BETSOL accelerates cloud transformation for enterprises across 17+ countries. BETSOL holds several engineering patents, and is recognized with industry awards. BETSOL maintains a net promoter score that is 2x the industry average.
BETSOL’s open source backup and recovery product line, Zmanda (Zmanda.com), delivers up to 50% savings in total cost of ownership (TCO) and delivers best-in-class performance.
BETSOL Global IT Services (BETSOL.com) builds and supports end-to-end enterprise solutions, reducing time-to-market for customers.
BETSOL offices are set against the vibrant backdrops of Broomfield, Colorado and Bangalore & Belagavi, India.
We take pride in being an employee-centric organization, offering comprehensive benefits and opportunities.
Learn more at betsol.com
Job Description
About the Role:
You will join the same technical pod that builds and operates our Kubernetes-based cloud platforms across Azure and GCP, working alongside our DevOps/DevSecOps engineers as a peer rather than a downstream verifier. Your focus is end-to-end solution validation of two platforms in parallel — Azure Local (Microsoft's hybrid on-prem stack, formerly Azure Stack HCI) and our FedRAMP-authorized GCP deployment — from initial deployment through Day-2 operations, upgrades, and decommission. You'll design and run tests across the full "ilities" spectrum, break things on purpose to prove resilience, and read Terraform/Ansible/Helm and observability data fluently enough to troubleshoot from platform down to code alongside engineering.
Responsibilities
Platform Validation
- Drive end-to-end validation for two Kubernetes-based platforms in parallel — Azure Local and the FedRAMP-authorized GCP deployment — covering initial deployment, Day-2 operations, upgrades, and decommission.
- Design, execute, and document solution-level test cases across reliability, availability, scalability, security, observability, performance, recoverability, and upgradability, written in Given/When/Then style that's readable by engineering, PM, SRE, and operations.
Kubernetes & Cloud Platform Testing
- Test and troubleshoot hands-on across kubectl, Helm, operators, CRDs, PVCs, StatefulSets, ArgoCD or Flux, network policies, and node affinity — able to debug from a pod down to the underlying node without a hand-hold.
- Validate Azure (AKS) and GCP (GKE) environments including Azure Monitor / GCP Cloud Monitoring, Key Vault / Cloud KMS, and Managed Identities / Workload Identity. Prior Azure Local experience isn't required; comfort learning a new hybrid on-prem topology quickly is.
Root Cause & Regression Testing
- Read Terraform, Ansible, and Helm charts fluently to understand what's deployed and why (writing IaC is not required).
- Reproduce customer or production incidents in the lab, trace the path from IaC through Kubernetes to cloud resources to root cause alongside engineering, and author the regression test that prevents recurrence.
Resilience & Chaos Testing
- Deliberately break and recover single nodes, pods, network paths, and storage (S2D pool degrade, PVC unbound, VPN tunnel drop) to verify SLOs hold under failure.
- Apply LitmusChaos, Chaos Mesh, or equivalent tooling where useful.
Observability-First Testing
- Verify system behavior by reading metrics, logs, and traces (Prometheus, Grafana, Loki, OpenTelemetry, Alertmanager) rather than asserting on HTTP status codes alone.
Security & FedRAMP-Aware Testing
- Test TLS termination, cert rotation, secret management (Cloud KMS / Secret Manager), RBAC boundaries, and container image scanning (Wiz, Prisma, or equivalent) on the GCP deployment.
- Verify audit logging into GCP Cloud Logging and generate evidence aligned to FedRAMP Moderate/High control families (AC, AU, SC, SI, IA); apply the same controls on the Azure Local side wherever they cross-map.
Collaboration
- Pair with engineers on design reviews, feature-freeze planning, and Day-2 playbooks as a peer — catching quality issues before code is written, not after.
Looking For
5+ years of QA / SDET experience, including at least 2 years hands-on with Kubernetes and 1+ year in a major public cloud (Azure or GCP).
Mandatory Skills
A. Technical
- Hands-on Kubernetes: kubectl, Helm, operators, CRDs, PVCs, StatefulSets, ArgoCD or Flux, network policies, node affinity.
- Azure and GCP fluency: AKS and GKE, Azure Monitor / GCP Cloud Monitoring, Key Vault / Cloud KMS, Managed Identities / Workload Identity.
- Ability to read Terraform, Ansible, and Helm charts to understand deployed infrastructure and trace issues to root cause.
- Test design and documentation across the full "ilities" spectrum (reliability, availability, scalability, security, observability, performance, recoverability, upgradability) in Given/When/Then style.
- Observability literacy: Prometheus, Grafana, Loki, OpenTelemetry, Alertmanager.
- Security and FedRAMP-aware testing: TLS/cert rotation, secret management, RBAC, container image scanning, audit-log verification, control-family evidence generation (AC, AU, SC, SI, IA).
B. Soft Skills
- Collaborates as a peer, not a downstream verifier — comfortable in design reviews and planning, not just execution.
- Clear, accountable communicator able to write test documentation for mixed technical and non-technical audiences.
- Strong problem-solving instincts and comfort operating across two platforms in parallel.
- Security mindset and outcome focus, measuring success by platform reliability and defect prevention.
Good to Have
- Zephyr or native Jira Test issuetype experience.
- LitmusChaos, Chaos Mesh, or equivalent chaos-engineering experience.
- Prior FedRAMP work, or experience in similar regulated environments (SOC 2, HIPAA, PCI DSS, DISA STIG).
- Experience with distributed data services: Kafka, Redis, MariaDB/Galera, MinIO or other S3-compatible object storage.
Qualifications
Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent professional experience.
