Skip to content

Platform Architect (AI/ML Infrastructure, GCP-focused)

  • Remote
    • Remote, São Paulo, Brazil
    • Remote, Buenos Aires, Argentina
    • Remote, Bolívar, Colombia
    • Remote, México, Mexico
    • remote, Arequipa, Peru
    +4 more
  • Software Development

Own GCP AI/ML infrastructure from model pipeline to reliable production service at scale.

My information

Fill out the information below

Upload your CV or resume file

Upload your cover letter

Questions

Please fill in additional questions

Have you completed the following level of education: Bachelor's Degree?
Do you have a degree in IT or Computer Science or equivalent experience?
Are you able to join the company in 2 weeks or less?
Do you have at least 5 years of experience in platform engineering, SRE, MLOps, or infrastructure, including operating production systems at scale?
Do you have at least 2 years of production experience operating Kubernetes clusters, including real failure modes and control-plane-level troubleshooting?
Do you have at least 2 years of production experience with ArgoCD or Flux, including GitOps and promotion workflows?
Do you have at least 3 years of experience designing or operating production infrastructure on GCP?
Do you have at least 2 years of production experience with Terraform, including complex state, reusable modules, multi-project configurations, and CI-driven plan/apply workflows?
Do you have at least 2 years of experience building and operating CI/CD pipelines, including experience with ML training or deployment pipelines?
Do you have at least 2 years of experience building infrastructure automation with Bash, Python, or Go, and do you actively use agentic coding tools?
In comparison to other professionals in the global tech market, where would you honestly rank yourself based on your technical expertise, experience, and achievements?
Are you currently the owner of a company?
Do you have hands-on experience deploying and operating ML or AI workloads in production, such as model serving, inference, or training infrastructure used by real users?
Do you have hands-on BigQuery experience in production—including partitioning and clustering, query cost/performance tuning, and dataset-level IAM?
Have you owned production reliability, including defining and measuring SLOs, incident response, post-mortems, and measurable reliability improvements?
Do you have experience with Dataflow, Pub/Sub, or Dataproc?
What is your preferred work location?