
Job Description
Ovii's Interpretation of the Role
Software Engineer on the Ray Data team building and scaling the Ray Datasets library. Focus on high‑performance distributed data processing, integration with ML workloads, and open‑source contributions.
Role Snapshot
- Build and optimize Ray Datasets library
- Improve performance at large scale
- Integrate with ML training pipelines
- Enhance testing and release processes
- Share knowledge via talks and tutorials
Must-Have Requirements
- Python
- Distributed Systems
- Algorithms & Data Structures
- System Design
- Data Processing frameworks (Spark, Dask)
- Ray
- algorithms, data structures, system design
- building scalable fault‑tolerant distributed systems
- data processing and database internals
Nice-to-Have Signals
- Streaming workloads experience
- Apache Arrow
- streaming workloads
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Develop high‑quality open‑source code for Ray Datasets
- Identify and implement architectural improvements to Ray core and the Datasets library
- Optimize performance of Ray Datasets for large‑scale workloads
- Integrate Ray Datasets with ML training and data source components
- Build and maintain stability and stress‑testing infrastructure
- Lead future work on streaming integrations such as Beam on Ray
- Differentiate data operations in Anyscale‑hosted Ray service
- Improve testing processes to ensure smooth releases
Good Fit If You Have
- Enjoy communicating technical work through talks, tutorials, or blog posts
- Comfortable working on performance‑critical, large‑scale systems
- Interest in open‑source ecosystems and community collaboration
Skills
- Python
- Distributed Systems
- Algorithms & Data Structures
- System Design
- Data Processing frameworks (Spark, Dask)
- Ray
- Apache Arrow (optional)
- Performance optimization
- Testing & CI