We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.
As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.
You'll Collaborate With
Compiler engineers, firmware & driver engineers, hardware architects, ML framework engineers and performance & validation teams
What You'll Own
The core runtime architecture that bridges the compiler and the hardware
Runtime execution engine performance and stability
Device memory management design and efficiency
Command submission and hardware interaction layer (in collaboration with firmware/driver teams)
Concurrency and scheduling model for model execution
Observability tooling (profiling, logging, tracing) within the runtime
Technical design decisions that ensure long-term scalability and maintainability
Minimum Qualifications
5–8+ years of experience in systems software, runtime systems, or performance-critical infrastructure
Strong proficiency in C/C++and Python
Solid understanding of fundamentals of operating systems, multi-threading, synchronization, memory management, and cache behavior
Experience working close to hardware (GPU, accelerator, driver, or firmware environments)
Familiarity with command queues, execution engines, and DMA-based systems
Understanding of computational graphs, tensor execution, and memory layouts
Experience profiling and optimizing latency and throughput in performance-sensitive systems
Preferred Qualifications
Experience building or architecting a runtime from scratch
Familiarity with MLIR, LLVM, or compiler-runtime interfaces
Experience with one or many from CUDA, ROCm, TensorRT, TVM, XLA, ONNX Runtime or similar runtimes
Experience integrating custom hardware backends into PyTorch, ONNX or OpenXLA
Knowledge of quantization and mixed-precision inference
Experience with multi-device or distributed execution
Background in AI accelerator or semiconductor environments
What Success Looks Like
You enable reliable end-to-end execution of compiled models on our NPU
Runtime overhead is minimal and performance targets are consistently met
Memory management and scheduling are efficient and scalable
The runtime architecture is robust, maintainable, and extensible
Compiler, firmware, and hardware teams can build confidently on top of your execution layer
The system supports real-world AI workloads with stability and observability
Join us in our mission to democratize AI compute — where your firmware expertise becomes the bedrock of tomorrow's AI breakthroughs.