JitLLM Chat Models
JitLLM (formerly GPULlama3.java) provides a Java-native implementation of LLMs that runs entirely in Java and executes automatically on GPUs via TornadoVM.
This extension allows Quarkus applications to use locally hosted LLMs (Llama3, Mistral, Qwen2.5, Deepseek-R1-Distill-Qwen, Qwen3, Qwen3.8, Gemma 4, Phi3, IBM Granite 3.2+, IBM Granite 4.0) for chat-based inference, leveraging GPU acceleration without requiring native CUDA code.
Here you can see a collection of demo Quarkus applications with jitllm.
Prerequisites
Java Version and TornadoVM SDK
The extension uses JitLLM 1.0.2, which is published in two lines, each paired with a TornadoVM 7.0.1 SDK for the same JDK line:
| JDK | JitLLM | TornadoVM SDK |
|---|---|---|
21 |
|
|
22 and newer (tested with 25) |
|
|
The build picks the JitLLM line that matches the JDK running Maven.
The JitLLM jar does not bundle TornadoVM; TornadoVM comes from the SDK at run time, through its tornado-argfile.
Install the JDK and the SDK with SDKMAN!:
|
Available TornadoVM backends
cuda, opencl, full (all backends). The extension names no backend itself; it runs on whichever one the selected SDK provides.
|
-
JDK 21:
# Java version sdk install java 21.0.2-open sdk use java 21.0.2-open # TornadoVM SDK sdk install tornadovm 7.0.1-jdk21-<backend> sdk use tornadovm 7.0.1-jdk21-<backend>
-
JDK 22 and newer (tested with JDK 25):
# Java version sdk install java 25.0.2-open sdk use java 25.0.2-open # TornadoVM SDK sdk install tornadovm 7.0.1-jdk22plus-<backend> sdk use tornadovm 7.0.1-jdk22plus-<backend>
To verify installation:
tornado --devices
tornado --version
The TornadoVM installation:
-
Sets the
TORNADOVM_HOMEenvironment variable to the TornadoVM SDK path. -
TORNADOVM_HOMEcontains thetornado-argfilewith all the JVM arguments required to enable TornadoVM. -
⚠️ The
tornado-argfileshould be used for building and running the Quarkus application (see section Building & Running the Quarkus Application).
Using JitLLM
To integrate the JitLLM chat model into your Quarkus application, add the following dependency:
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-jitllm</artifactId>
<version>1.14.1</version>
</dependency>
Even better, if you use the Quarkus platform BOM (default for projects generated), add the Quarkus Langchain4J BOM and all dependency versions will align:
<dependencyManagement>
<dependencies>
<dependency>
<groupId>${quarkus.platform.group-id}</groupId>
<artifactId>${quarkus.platform.artifact-id}</artifactId>
<version>${quarkus.platform.version}</version>
<type>pom</type>
<scope>import</scope>
</dependency>
<dependency>
<groupId>${quarkus.platform.group-id}</groupId>
<artifactId>quarkus-langchain4j-bom</artifactId> (1)
<version>${quarkus.platform.version}</version> (2)
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-jitllm</artifactId>
(3)
</dependency>
</dependencies>
| 1 | In your dependencyManagement section, add the quarkus-langchain4j-bom |
| 2 | Inherit the version from your platform version |
| 3 | Voilà, no need for version alignment anymore |
| If no other LLM extension is configured, AI Services will automatically use the GPU-accelerated JitLLM model! |
Sample implementation for ChatModel:
@Path("chat")
public class ChatLanguageModelResource {
private final ChatModel chatModel;
public ChatLanguageModelResource(ChatModel chatModel) {
this.chatModel = chatModel;
}
@GET
@Path("blocking")
public String blocking() {
return chatModel.chat("When was the nobel prize for economics first awarded?");
}
}
Send requests to blocking endpoint:
curl http://localhost:8080/chat/blocking
Sample implementation for StreamingChatModel:
@Path("chat")
public class ChatLanguageModelResource {
private final StreamingChatModel streamingChatModel;
public ChatLanguageModelResource(StreamingChatModel streamingChatModel) {
this.streamingChatModel = streamingChatModel;
}
@GET
@Path("streaming")
@RestStreamElementType(MediaType.TEXT_PLAIN)
public Multi<String> streaming() {
return Multi.createFrom().emitter(emitter -> {
streamingChatModel.chat("When was the nobel prize for economics first awarded?",
new StreamingChatResponseHandler() {
@Override
public void onPartialResponse(String token) {
emitter.emit(token);
}
@Override
public void onError(Throwable error) {
emitter.fail(error);
}
@Override
public void onCompleteResponse(ChatResponse completeResponse) {
emitter.complete();
}
});
});
}
}
Send requests to streaming endpoint:
curl http://localhost:8080/chat/streaming
Configure JitLLM
The JitLLM extension can be configured via standard Quarkus properties:
# Enable JitLLM integration
quarkus.langchain4j.jitllm.enable-integration=true
# Select the default model
quarkus.langchain4j.jitllm.chat-model.model-name=unsloth/Llama-3.2-1B-Instruct-GGUF
quarkus.langchain4j.jitllm.chat-model.quantization=F16
quarkus.langchain4j.jitllm.chat-model.temperature=0.7
quarkus.langchain4j.jitllm.chat-model.max-tokens=1024
Model files are automatically downloaded from Beehive Lab HuggingFace if not available locally.
To use a GGUF file you already have, point quarkus.langchain4j.jitllm.models-path at a directory laid out the way the cache is: <owner>_<name>/<name>-<quantization>.gguf and an empty .finished marker next to it. For example, local_Qwen3-4B/Qwen3-4B-Q8_0.gguf is selected with model-name=local/Qwen3-4B and quantization=Q8_0. The model name, quantization and models path are fixed at build time.
|
Building & Running the Quarkus Application
Dev Mode
To run your Quarkus application in dev mode with TornadoVM:
-
Ensure your
pom.xmlcontains thequarkus-langchain4j-jitllmdependency (shown earlier). -
Add the TornadoVM argfile as a Maven property:
<properties> <tornado.argfile>${env.TORNADOVM_HOME}/tornado-argfile</tornado.argfile> </properties> -
Pass the argfile to the JVM in the plugin configuration for dev mode:
<plugin> <groupId>io.quarkus</groupId> <artifactId>quarkus-maven-plugin</artifactId> <configuration> <jvmArgs>@${tornado.argfile}</jvmArgs> </configuration> </plugin> -
Launch dev mode explicitly:
mvn quarkus:dev
Production Mode
To build and run your application in production mode:
-
Build the Quarkus application:
mvn clean package -
Run the generated jar with the TornadoVM argfile:
java @$TORNADOVM_HOME/tornado-argfile --add-modules jdk.incubator.vector \ -jar target/quarkus-app/quarkus-run.jar
Ensure TORNADOVM_HOME is set and points at the SDK matching your JDK. The tornado-argfile does not add jdk.incubator.vector, which the engine needs, so pass --add-modules jdk.incubator.vector yourself.
|
The model’s device-memory budget is set by quarkus.langchain4j.jitllm.chat-model.device-memory (default 4GB). The extension sets tornado.device.memory from it at model initialization, so a bare -Dtornado.device.memory on the command line has no effect.
|
Supported Models and Quantizations
The following models have been tested with JitLLM and can be found in Beehive Lab’s HuggingFace Collections.
|
Quantization format names are model-dependent. Some models require |
| Model | Quantizations | Model Identifier (Hugging Face) |
|---|---|---|
Llama 3.2 1B |
|
|
Llama 3.2 3B |
|
|
Llama 3.1 8B |
|
|
Mistral 7B |
|
|
Qwen2.5 0.5B |
|
|
Qwen2.5 1.5B |
|
|
DeepSeek-R1 Distill Qwen 1.5B |
|
|
DeepSeek-R1 Distill Qwen 7B |
|
|
Qwen3 0.6B |
|
|
Qwen3 1.7B |
|
|
Qwen3 4B |
|
|
Qwen3 8B |
|
|
Phi 3 Mini 4k |
|
|
Phi 3 Mini 4k |
|
|
Phi 3 Mini 128k |
|
|
Phi 3.1 Mini 128k |
|
|
Granite 3.2 2B |
|
|
Granite 3.2 8B |
|
|
Granite 3.3 2B |
|
|
Granite 3.3 8B |
|
|
Granite 4.0 1B |
|
|
Each entry corresponds to a GGUF model tested to run on TornadoVM via jitllm.
Configuration Reference
Configuration property fixed at build time - All other configuration properties are overridable at runtime
Configuration property |
Type |
Default |
|---|---|---|
Determines whether the necessary JitLLM models are downloaded and included in the jar at build time. Currently, this option is only valid for Environment variable: |
boolean |
|
Whether the model should be enabled Environment variable: |
boolean |
|
Model name to use Environment variable: |
string |
|
Quantization of the model to use Environment variable: |
string |
|
Location on the file-system which serves as a cache for the models Environment variable: |
path |
|
What sampling temperature to use, between 0.0 and 1.0. Environment variable: |
double |
|
What sampling topP to use, between 0.0 and 1.0. Environment variable: |
double |
|
What seed value to use. Environment variable: |
int |
|
The maximum number of tokens to generate in the completion. Environment variable: |
int |
|
Whether to use the prefill/decode inference path, which processes the prompt in a dedicated prefill phase before single-token decoding. Combined with a Off by default. Batched prefill is default-off in the engine on every backend, and enabling it is an opt-in that needs its own performance evidence for the model and device in question. Environment variable: |
boolean |
|
Number of prompt tokens processed per batch during the prefill phase. Only used when Environment variable: |
int |
|
Whether to enable the model’s thinking/reasoning phase. Only models with a thinking mode (e.g. Qwen3) honor this; for other models it has no effect. When Environment variable: |
boolean |
|
Amount of GPU memory TornadoVM is allowed to allocate e.g. Note: this maps to the JVM-global engine flag Environment variable: |
string |
|
Whether to enable the integration. Set to Environment variable: |
boolean |
|
Whether JitLLM should log requests Environment variable: |
boolean |
|
Whether JitLLM client should log responses Environment variable: |
boolean |
|
Type |
Default |
|
Model name to use Environment variable: |
string |
|
Quantization of the model to use Environment variable: |
string |
|
What sampling temperature to use, between 0.0 and 1.0. Environment variable: |
double |
|
What sampling topP to use, between 0.0 and 1.0. Environment variable: |
double |
|
What seed value to use. Environment variable: |
int |
|
The maximum number of tokens to generate in the completion. Environment variable: |
int |
|
Whether to use the prefill/decode inference path, which processes the prompt in a dedicated prefill phase before single-token decoding. Combined with a Off by default. Batched prefill is default-off in the engine on every backend, and enabling it is an opt-in that needs its own performance evidence for the model and device in question. Environment variable: |
boolean |
|
Number of prompt tokens processed per batch during the prefill phase. Only used when Environment variable: |
int |
|
Whether to enable the model’s thinking/reasoning phase. Only models with a thinking mode (e.g. Qwen3) honor this; for other models it has no effect. When Environment variable: |
boolean |
|
Amount of GPU memory TornadoVM is allowed to allocate e.g. Note: this maps to the JVM-global engine flag Environment variable: |
string |
|
Whether to enable the integration. Set to Environment variable: |
boolean |
|
Whether JitLLM should log requests Environment variable: |
boolean |
|
Whether JitLLM client should log responses Environment variable: |
boolean |
|
Limitations
-
TornadoVM currently does not support GraalVM Native Image builds.
-
Ensure that the TornadoVM environment (
TORNADOVM_HOMEandtornado-argfile) is properly set before running Quarkus. -
Only Java 21 or newer versions are supported.
-
Tool calling is supported for the Llama 3.x, Qwen2.5, Qwen3, Qwen3.8, Gemma 4 and Granite 3.2/4.0 families. It is not supported for Phi-3, DeepSeek-R1-Distill, Qwen MoE and Mistral/Devstral: a request that carries tools for one of these models is refused up front with an
UnsupportedFeatureException, rather than sent to a model that cannot call them. -
Forced tool choice (
ToolChoice.REQUIRED) and a JSON response format are not supported; requests that set either are rejected. -
Gemma 4 and Qwen3.8 run on both the CUDA and the OpenCL backends, as do the other families.