The English-language job board for Spain
Senior Cloud Infrastructure and DevOps Solutions Architect
Remote from Spain·Added 15 days ago
Overview
Job details
Fully remote
Per the ad.
Requirements
Have the right to work in Spain
NVIDIA doesn't mention sponsorship in the ad.
Have 8+ years of experience
Senior-level role.
University degree
“BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.” — per the ad.
Requirements
What we're looking for
8+ years in managing scalable cloud environments and automation engineering roles.
Cloud, HPC & GPU Expertise: Proven understanding of networking fundamentals and data centre architectures, with hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure-deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high-speed interconnects.
Kubernetes & AI/ML Workloads: Extensive background with Kubernetes for container orchestration, resource scheduling and scaling in GPU-accelerated and HPC environments, including scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised (e.g. KubeVirt) multi-tenant estates.
Linux & Storage Systems: Deep knowledge of Linux (RedHat, Ubuntu), OS-level security, and protocols. Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and emerging Kubernetes storage technologies.
Automation, GitOps & Observability: Proficiency in Python and Bash scripting, configuration management and Infrastructure-as-Code tools (e.g. Ansible, Terraform), GitOps-based cluster lifecycle and upgrade management for large fleets, and observability stacks (Grafana, Loki, Prometheus) for monitoring, logging and building fault-tolerant systems.
Fleet Reliability & Customer Engagement: Demonstrated ability to measure and improve MTBI and job goodput on large GPU clusters-fault detection, drain and remediation workflows, SLO/error-budget definition and post-incident review-combined with a strong consultative background leading architectural reviews and presenting to executive stakeholders.
Nice to have
Knowledge of CI/CD pipelines and container-based microservices architectures for software deployment and automation.
Experience with the NVIDIA GPU and Network Operators for automated GPU and network resource lifecycle management in Kubernetes, and with NVIDIA Base Command Manager (BCM) for provisioning, managing and monitoring GPU clusters at scale.
Familiarity with GPU health and fleet telemetry tooling-DCGM and XID diagnostics, node-level health agents, and fleet-wide reliability intelligence.
Expertise in AI-native scheduling and inference frameworks on Kubernetes (e.g. KAI, Grove, Dynamo, NVIDIA Cloud Functions).
Background with RDMA-based fabrics (InfiniBand or RoCE) in HPC or AI environments. Exposure to Cumulus Linux, SONiC or Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations (NMX-C / NMX-M) on NVL72-class systems is a strong plus.
Education & certifications
BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
The role
NVIDIA is looking for a Senior Cloud Infrastructure and DevOps Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial organizations around the world are using NVIDIA products to redefine deep learning and data analytics, and to power next-generation data centers. Join the team building and advising on many of the largest and fastest AI/HPC systems in the world!
What you'll do
We are looking for someone who combines deep technical expertise with strong consulting and communication skills. This role will engage directly with customers, partners, and cross-functional teams to assess, architect, and guide the implementation of large-scale infrastructure projects. The scope spans system architecture, Kubernetes-based platforms, and automation-serving as both a trusted advisor and a hands-on technical leader. You will sit at the centre of NVIDIA's Cloud Partner (NCP) operating model, covering the full Day 1 to Day 2 lifecycle: taking a GPU cluster from hardware handover, through full-solution validation, to a production-stable platform running at maximum goodput. NCP estates are open-source-first and heterogeneous-upstream Kubernetes, KubeVirt, Slurm, Prometheus/Grafana, Cumulus/SONiC and a long tail of ISV software-so this role is deliberately tool-agnostic: you will meet each partner on the stack they actually run rather than on a single proprietary product.
Own full-solution validation on the partner software stack-the layer above hardware validation-including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in against agreed MTBI and goodput targets.
Minimise the time from cluster handover to first production workload, working across hardware bring-up, managed-service intake and the partner's own operations teams to remove duplicated validation and handover friction.
Own Day 2 production stability at fleet scale: monitoring, logging and workload orchestration, fault detection and remediation, preventive maintenance, and proactive firmware and field-notice rollout campaigns.
Assess customer environments and operate heterogeneous open platforms-upstream Kubernetes, KubeVirt, Slurm and GPU-aware schedulers-integrated with enterprise-grade networking and storage, and enable third-party ISV workloads on top of them.
Provide consultative guidance and hands-on troubleshooting across the full stack-bare metal, operating system, software stack, container platform, networking and storage-and support R&D, POCs and POVs validating new features, architectures and upgrade approaches.
Act as the technical leader for assigned accounts: run structured knowledge transfer and enablement, and produce runbooks, onboarding materials and best-practice guides so partner teams can operate advanced configurations independently.
Hiring process
Search and apply
Phone interviews
Phone call
Full-time candidates typically start with phone interviews.
Virtual or in-person interviews
With hiring manager, team members and employees from other groups · 30–60 min
One-on-one, small-group or panel interviews with the hiring manager, team members and people from other groups; an onsite in-person interview at an office is required before an offer.
Insider Chat
With a community resource group member · 15 min · Not every role
An optional 15-minute conversation with a community resource group member that does not influence the hiring decision.
Decision and offer
- Applicants selected to move forward hear from recruiting within a couple of weeks.
- Most candidates have a decision within weeks of their first interview.
- Technical candidates may complete coding exercises on HackerRank.
About NVIDIA
NVIDIA is a global technology company that designs and manufactures graphics processing units (GPUs) and system-on-a-chip units for gaming, professional visualization, data centers, and artificial intelligence. Founded in 1993, NVIDIA has grown into a multinational corporation with a market value exceeding $1 trillion, employing over 26,000 people worldwide. The company is a household name in the tech industry, powering everything from gaming consoles to the world's most advanced AI supercomputers.
In Spain, NVIDIA has a significant presence with offices in Madrid and Barcelona, focusing on sales, developer relations, and solutions architecture for the Southern European and EMEA regions. The company is known for a high-performance, innovation-driven culture that attracts top engineering talent globally. For international professionals, NVIDIA offers competitive compensation, a strong brand name, and the opportunity to work on cutting-edge AI and HPC technologies, making it a highly attractive employer for those looking to relocate to Spain.
- Industry
- Technology
- Founded
- 1993
- Employees
- 26,000–30,000
- Headquarters
- Santa Clara, USA
- In Spain
- Madrid, Barcelona
- Website
- nvidia.com
Good to know if you are moving
- NVIDIA has offices in Madrid and Barcelona, with a strong focus on Southern Europe and EMEA markets.
- The company is a global leader in AI computing, offering engineers the chance to work on cutting-edge technologies like CUDA, cuDNN, and the NVIDIA AI Enterprise platform.
- NVIDIA's work culture is known for being fast-paced, innovative, and performance-driven, with a strong emphasis on technical excellence.
- The company offers competitive compensation and benefits packages, and is consistently rated as one of the best places to work in the tech industry.
- For international hires, NVIDIA provides visa sponsorship and relocation support, making it a viable option for professionals looking to move to Spain.
More jobs like this
or browse DevOps & SRE·Senior·Remote·Technology








