Systems Research Engineer Intern - GPU Programming (Winter 2027) at Together AI in San Francisco
- Company: Together AI
- Location: San Francisco
- Posted: Sep 18, 2026
- Type: Internship
- Salary: $58 to $70 an hour
Overview
As a Systems Research Engineer Intern specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. Working closely with the modeling and algorithm team, you will co-design GPU kernels and model architecture t…
Job description
- As a Systems Research Engineer Intern specialized in GPU Programming, you will play a crucial role in developing and optimizing GPU-accelerated kernels and algorithms for ML/AI applications. Working closely with the modeling and algorithm team, you will co-design GPU kernels and model architecture to enhance the performance and efficiency of our AI systems. Collaborating with the hardware and software teams, you will contribute to the co-design of efficient GPU architectures and programming models, leveraging your expertise in GPU programming and parallel computing. Your research skills will be vital in staying up-to-date with the latest advancements in GPU programming techniques, ensuring that our AI infrastructure remains at the forefront of innovation.
- This internship is based on-site at our San Francisco HQ, running through the Winter term from January to April.
Responsibilities
- Optimize and fine-tune GPU code to achieve better performance and scalability
- Collaborate with cross-functional teams to integrate GPU-accelerated solutions into existing software systems
- Stay up-to-date with the latest advancements in GPU programming techniques and technologies
Requirements
- Strong background in GPU programming and parallel computing, such as CUDA and/or Triton.
- Knowledge of ML/AI applications and models
- Knowledge of performance profiling and optimization tools for GPU programming
- Excellent problem-solving and analytical skills
Skills
Required
- GPU programming
- Parallel computing
- CUDA
- Triton
- ML/AI applications and models
Benefits
- We offer competitive compensation, housing stipends, and other competitive benefits. The estimated US hourly rate for this role is $58 to $70 an hour. Our hourly rates are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
About Together AI
Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.