Staff Data Engineer (Audio/ML)
Nicasio, CA, USA Hybrid
Salary not listedJobFig found this opening at its original source and checks that it remains available.
About the role
The Skywalker Sound Development Group is seeking an experienced Data Engineer with a focus on Audio/ML to specialize in the creation, management, and optimization of data pipelines to support cutting-edge AI/ML research. This is a critical role in preparing high-quality datasets for the training, retraining, and evaluation of machine learning models tailored to immersive and multichannel audio applications. As a Data Engineer (Audio/ML), you will focus on developing robust pipelines for processing complex media datasets, enabling AI/ML researchers to build transformative solutions for speech processing, style transfer, and source separation. Your work will directly contribute to creating innovative soundtrack workflows for global media production. This role is considered Hybrid, which means the employee will work 2-3 days onsite at our Nicasio, CA office and occasionally from home.
What you'll bring
- 8+years of experience in data engineering or data science with a focus on building pipelines for AI/ML applications.
- Proficiency in Python, with expertise in data manipulation libraries such as Pandas, NumPy, and PyTorch’s data utilities.
- Hands-on experience with audio processing libraries and tools (e.g., Librosa, FFmpeg, SoX) for handling complex audio formats.
- Familiarity with scalable pipeline tools like GitLab, Apache Spark, Airflow, or Luigi, and experience with containerized workflows (Docker, Kubernetes).
- Strong understanding of data pipeline requirements for model training, retraining, and evaluation in iterative research workflows.
- Experience with immersive and multichannel audio formats.
- Knowledge of cloud-based platforms and tools for storage and processing, such as AWS S3, Redshift, or Google BigQuery.
- Strong problem-solving skills, with a proactive mindset for addressing evolving data challenges.
- Master’s Degree with preference for PhD in Data Engineering/Science, Computer Science, Signal Processing, or a related field.
- Experience integrating data pipelines with AI/ML workflows, including active learning and model retraining.
- Familiarity with audio-specific datasets and metadata management strategies.
- Knowledge of machine learning principles and how data quality impacts model performance.