Get up and running with Llama 3, Mistral, Gemma, and other large language models.
GoMIT Licenseactive
17 projectsLLM Inference
Get up and running with Llama 3, Mistral, Gemma, and other large language models.
GoMIT Licenseactive
A high-throughput and memory-efficient inference and serving engine for LLMs.
PythonApache License 2.0active
A distributed multi-model LLM serving system with web UI and OpenAI-compatible RESTful APIs.
PythonApache License 2.0steady
SGLang is a fast serving framework for large language models and vision language models.
PythonApache License 2.0active
Nvidia Framework for LLM Inference
PythonOtheractive
Fine-tune, serve, deploy, and monitor any open-source LLMs in production. Used in production at BentoML for LLMs-based applications.
PythonApache License 2.0active
Run LLMs and batch jobs on any cloud. Get maximum cost savings, highest GPU availability, and managed execution -- all with a simple interface.
PythonApache License 2.0active
A high-throughput and low-latency inference and serving framework for LLMs and VLs
PythonApache License 2.0active
Blazingly fast LLM inference.
RustMIT Licenseactive
Efficient Triton Kernels for LLM Training.
PythonBSD 2-Clause "Simplified" Licenseactive
NVIDIA Framework for LLM Inference(Transitioned to TensorRT-LLM)
C++Apache License 2.0slowing
Inference for text-embeddings in Rust, HFOIL Licence.
RustApache License 2.0active
A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
PythonMIT Licenseslowing
Inference for text-embeddings in Python
PythonMIT Licensesteady
To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
PythonMIT Licensesteady
Reference: llm-inference-solutions
MIT Licenseslowing
Easily deploy any LLM on a VM with minimal configuration, using Ansible.
slowing