C

Distributed Inference Engineer — LLM Engine

Remote
Apply
AI Summary

Design, build, and optimize distributed systems for large language model inference at scale. Develop core inference infrastructure including model sharding, load balancing, and low-latency networking components. Collaborate with engineering teams to improve performance, reliability, and cost efficiency across production environments.

Key Highlights
Founding team role shaping core technical architecture and engineering culture
Full-time remote position with asynchronous collaboration and code reviews
Requires systems programming expertise in Rust, C++, or Go for production-grade services
Focus on distributed inference optimization, GPU utilization, and scalable microservices
Key Responsibilities
Design, build, and optimize distributed systems that serve large language models at scale
Develop and maintain core inference infrastructure including model sharding, load balancing, caching, and low-latency networking components
Profile performance, reduce inference costs, improve throughput, and enhance reliability across diverse hardware configurations
Write production-grade code, implement monitoring and observability, and debug complex distributed issues
Contribute to architectural decisions for the LLM engine
Technical Skills Required
Rust C++ Go
Benefits & Perks
Remote work

Job Description


Company Description Cognition Commons is an early-stage organization focused on building advanced large language model (LLM) infrastructure and tools that make AI systems more accessible, reliable, and efficient. The team is committed to pushing the boundaries of distributed inference, model deployment, and scalable AI operations. As part of the founding team, individuals have the opportunity to shape core technical architecture, engineering culture, and long-term product direction. Cognition Commons values collaborative problem-solving, rigorous engineering standards, and transparent decision-making. The company supports remote work and aims to attract builders who are excited about creating foundational AI technology from the ground up.
Role Description As a Distributed Inference Engineer — LLM engine on the founding team, you will design, build, and optimize distributed systems that serve large language models at scale. You will develop and maintain core inference infrastructure, including model sharding, load balancing, caching, and low-latency networking components for remote production environments. You will work closely with other engineers to profile performance, reduce inference costs, improve throughput, and enhance reliability across diverse hardware configurations. Your day-to-day responsibilities will include writing production-grade code, implementing monitoring and observability, debugging complex distributed issues, and contributing to architectural decisions for the LLM engine. This is a full-time remote role that involves close collaboration via asynchronous communication, code reviews, and regular technical design discussions.
Qualifications
  • Strong software engineering skills in systems programming languages (e.g., Rust, C++, Go) and experience building high-performance, production-grade services.
  • Experience with distributed systems concepts such as concurrency, fault tolerance, consensus, and scalable microservices architectures.
  • Familiarity with machine learning model serving, GPU/accelerator utilization, and frameworks or runtimes used for LLM inference.
  • Comfort with cloud infrastructure, containerization, and orchestration tools (e.g., Kubernetes, Docker, major cloud providers).
  • Proficiency with monitoring, logging, and performance profiling tools to diagnose and resolve bottlenecks in distributed environments.
  • Ability to write clear technical documentation, participate in rigorous code reviews, and communicate effectively in remote, distributed teams.
  • Prior experience in early-stage startups or founding engineering teams, with a willingness to take ownership and operate in a fast-changing environment.
  • Bachelor’s or advanced degree in Computer Science, Engineering, or a related field, or equivalent practical experience in systems and infrastructure engineering.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

claven

Serbia
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Entry level

roosevelt marketing outstaffin...

Serbia

Senior Kernel Developer - KernelCare Team

Programming
1w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

CloudLinux

Serbia

Subscribe our newsletter

New Things Will Always Update Regularly