Job Opportunity Posted yesterday

AI model optimization & acceleration Engineer

L&T Technology Services
Bangalore

Job Description

Project Details

AI model optimization & acceleration


Job Description

Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU).

Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive.


Experience : 4-10 Years


Job Responsibilities / Day-to-Day Activities


Qualifications & Experiences:

Key Responsibilities

• Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech

• Port models across frameworks (e.g., PyTorch → ONNX → runtimes)

• Deploy on hardware accelerators (GPU/NPU) and optimize performance

• Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion)

• Apply quantization and model compression (FP32 → lower precision)

• Profile and debug system and model performance


Required Skills

• Strong in PyTorch (or similar), ONNX (or equivalent)

• Proficient in Python and C++

• Experience with GPU/hardware acceleration (CUDA/ROCm or similar)

• Solid understanding of deep learning models (transformers, CNNs)

• Knowledge of optimization, quantization, and performance tuning


Good to Have

• Edge AI or embedded deployment

• Generative or multi-modal AI systems

• Distributed inference or streaming pipelines

About this job listing
This job opportunity is provided through our external job listing network. MyJobAlerts helps you discover job opportunities and redirects you to the original listing to apply.