跳到正文
Hugging Face Blog·· 2025-01-16精选AI 评分61

Hugging Face 为 TGI 引入多后端架构,支持 TensorRT-LLM 与 vLLM 等引擎

Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference

AI 导读

Hugging Face 宣布为 Text Generation Inference(TGI)引入多后端架构,允许将其作为统一前端层接入 TensorRT-LLM、vLLM、llama.cpp 等多种底层推理引擎。

推荐理由

原文阐述了将 TGI 作为统一前端接入多种推理引擎的架构设计,读者可以了解跨硬件推理服务的解耦与选型方案。

来源:Hugging Face Blog · huggingface.co