ONHUMAN is an AI robotics company developing autonomous general-purpose humanoid robots. Our goal is to build embodied AI systems that can perceive, reason, and act in the real world. ONHUMAN is headquartered in Madrid, Spain, and this role requires 5 days/week in-office collaboration.
Team & Role Overview
Our ONVULC team is responsible for developing the core AI systems that power humanoid autonomy. We are looking for an AI Engineer, Video Training to design and advance the large-scale video pretraining infrastructure, video generation pipelines, and predictive world models that allow our robots to visually simulate and reason about the physical world.
This role focuses on developing new architectures for video representation and generation—spanning massive video tokenization, spatial-temporal diffusion models, and video-to-video future prediction frameworks that directly scale robot intelligence.
Responsibilities
Video Model Architecture: Design and develop generative video model architectures (e.g., video diffusion transformers) capable of high-fidelity future prediction and visual world modeling.
Large-Scale Data Engineering: Build scalable pipelines to ingest, curate, tag, and filter billions of frames of real-world and synthetic video data for self-supervised pretraining.
Spatial-Temporal Learning: Advance video modeling approaches across time and space, including novel spatial-temporal attention mechanisms, video tokenization (VQ-VAE/ViViT variants), and long-context video understanding.
Capability & Scalability Optimization: Improve training throughput, model stability, and scaling laws for multi-modal generative video systems distributed across thousands of GPUs.
End-to-End Model Lifecycle: Work across the model training lifecycle, from initial architectural research and tokenization exploration to distributed pretraining and edge inference optimization.
Cross-Functional Fusion: Collaborate closely with modeling, pretraining, generative AI, and RL teams to feed high-quality video representations and predicted world states into the robot autonomy stack.
Experimental Evaluation: Design experiments and diagnostic frameworks to evaluate video coherence, geometric consistency, physical plausibility, and downstream performance on robotic tasks.
Paradigm Development: Contribute to the development of next-generation video pretraining and world modeling paradigms tailored explicitly for physical, embodied AI intelligence.
Requirements
Generative Video/Vision Experience: Experience designing and training large-scale deep learning models for video generation, future frame prediction, or advanced visual representation systems.
Modern Generative Architectures: Deep understanding of modern generative paradigms, including Diffusion Models, Autoregressive Transformers, Flow Matching, and related spatial-temporal architectures.
Distributed Training Proficiency: Proven experience training frontier models on large GPU clusters utilizing distributed training frameworks (e.g., Megatron-LM, DeepSpeed, FSDP, PyTorch Distributed).
Technical Proficiency: Mastery of Python and deep learning frameworks such as PyTorch, along with custom CUDA kernel optimization or high-performance I/O pipelines.
Experimental Rigor: Exceptionally strong technical rigor and the data-driven mindset needed to diagnose, stabilize, and optimize loss curves on high-compute model training runs.
Software Engineering Skills: Solid software engineering skills with the ability to build reliable, reusable, and maintainable data processing and model infrastructure.
Autonomy & Ambiguity: Ability to operate independently and drive ambiguous, multi-million-parameter technical challenges from research to deployment.
Bonus Qualifications
World Models for Robotics: Background in learning world models, predictive dynamics, or self-supervised video representation specifically applied to physical control loops.
Frontier Video Generative Labs: Experience working on cutting-edge video generation platforms (e.g., Sora, Gen-3, Movie Gen, Lumiere) at companies like OpenAI, Google DeepMind, Runway, Meta, or Anthropic.
Synthetic Data Generation: Experience combining video models with 3D simulators (e.g., Unreal Engine, MuJoCo, Isaac Sim) to create high-fidelity synthetic training loops.
Publication Record: A strong publication record in machine learning, computer vision, or generative modeling conferences (e.g., CVPR, ICCV, NeurIPS, ICLR).
Compensation
The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
AI - ONVULC Neural Network





Benefits
Madrid, Spain
€80,000 Anually
Equity Plan
Flex Schedule
Growth Budget
Modern Tools