Built a complete optimization pipeline to solve the problem of high memory consumption and slow inference of deep learning models. Adopted knowledge distillation, quantization, and pruning for model compression. Combined CUDA and torch.compile for inference acceleration. Created automated evaluation scripts to compare latency, memory usage, and accuracy before and after optimization, lowering hardware deployment costs for edge-side model deployment.
需求方专属客服,免费梳理匹配
添加客服微信,免费为您安排与该工程师直接沟通
长按二维码添加客服微信
如当前档期不合,也可免费为您推荐相似案例作者