Jobs / engineering
Intern - ML Inference Performance Engineer
AXELERA AI · Eindhoven
What you will do
Benchmarking & Tooling: Develop a thorough understanding of internal benchmarking tools covering throughput, latency, power, and accuracy across device-level, host-transaction, and end-to-end pipeline scenarios. Improve existing tooling, define reproducible procedures, and establish a standardised results format for rigorous cross-platform comparisons. Maintain a dedicated dashboard for performance visualisations.
Platform Evaluation: Research and evaluate AI accelerator products from various vendors, gaining hands-on experience with their SDKs, toolchains, flexibility, and limitations through a structured evaluation process. Track model support across platforms to identify strengths, gaps, and areas for improvement.
Pipeline Analysis: Characterise full inference pipelines, capturing host-device transaction overhead and end-to-end performance metrics. Ensure equivalent pipeline configurations across platforms using frameworks such as GStreamer to maintain methodological consistency.
Lab & Infrastructure: Set up and maintain lab hosts across multiple hardware platforms and support the onboarding of new evaluation hardware.
Reporting: Synthesise findings into clear, structured reports that directly inform engineering and roadmap decisions.