Skip to main content
I

Inferact

Inferact builds AI inference infrastructure to make model serving as simple as spinning up a serverless database; its founders created and maintain vLLM.

Inferact builds AI inference infrastructure intended to narrow the gap between capable AI models and the systems that serve them. The company's stated goal is to make serving AI as simple as spinning up a serverless database, thereby reducing the cost and latency of inference. Its technical work spans LLM inference, model serving, hardware accelerator optimisation, serverless infrastructure, and support for mixture-of-experts, multimodal and agentic AI.

The company was founded by the creators and core maintainers of vLLM, the open-source LLM inference engine they built and continue to maintain. vLLM supports more than 500 model architectures and runs on more than 200 accelerator types at global scale, with over 2,000 contributors to the project. Alongside this open-source work, Inferact is developing its own inference infrastructure product.

Inferact operates with a global footprint and is rooted in open-source, community-driven development. Its stated mission is to accelerate AI progress and to make AI infrastructure accessible beyond organisations able to build their own custom stacks. The technical domains and open-source community central to the company are the substance of any engineering or technical-marketing role there.

Open jobs

No open jobs right now.