JitLLM Chat Models

JitLLM (formerly GPULlama3.java) provides a Java-native implementation of LLMs that runs entirely in Java and executes automatically on GPUs via TornadoVM.

This extension allows Quarkus applications to use locally hosted LLMs (Llama3, Mistral, Qwen2.5, Deepseek-R1-Distill-Qwen, Qwen3, Qwen3.8, Gemma 4, Phi3, IBM Granite 3.2+, IBM Granite 4.0) for chat-based inference, leveraging GPU acceleration without requiring native CUDA code.

Here you can see a collection of demo Quarkus applications with jitllm.

Prerequisites

Java Version and TornadoVM SDK

The extension uses JitLLM 1.0.2, which is published in two lines, each paired with a TornadoVM 7.0.1 SDK for the same JDK line:

JDK JitLLM TornadoVM SDK

21

io.github.beehive-lab:jitllm:1.0.2-jdk21

7.0.1-jdk21-<backend>

22 and newer (tested with 25)

io.github.beehive-lab:jitllm:1.0.2-jdk22plus

7.0.1-jdk22plus-<backend>

The build picks the JitLLM line that matches the JDK running Maven. The JitLLM jar does not bundle TornadoVM; TornadoVM comes from the SDK at run time, through its tornado-argfile. Install the JDK and the SDK with SDKMAN!:

Available TornadoVM backends
cuda, opencl, full (all backends). The extension names no backend itself; it runs on whichever one the selected SDK provides.
  • JDK 21:

# Java version
sdk install java 21.0.2-open
sdk use java 21.0.2-open

# TornadoVM SDK
sdk install tornadovm 7.0.1-jdk21-<backend>
sdk use tornadovm 7.0.1-jdk21-<backend>
  • JDK 22 and newer (tested with JDK 25):

# Java version
sdk install java 25.0.2-open
sdk use java 25.0.2-open

# TornadoVM SDK
sdk install tornadovm 7.0.1-jdk22plus-<backend>
sdk use tornadovm 7.0.1-jdk22plus-<backend>

To verify installation:

tornado --devices
tornado --version

The TornadoVM installation:

  • Sets the TORNADOVM_HOME environment variable to the TornadoVM SDK path.

  • TORNADOVM_HOME contains the tornado-argfile with all the JVM arguments required to enable TornadoVM.

  • ⚠️ The tornado-argfile should be used for building and running the Quarkus application (see section Building & Running the Quarkus Application).

Using JitLLM

To integrate the JitLLM chat model into your Quarkus application, add the following dependency:

<dependency>
  <groupId>io.quarkiverse.langchain4j</groupId>
  <artifactId>quarkus-langchain4j-jitllm</artifactId>
  <version>1.14.1</version>
</dependency>

Even better, if you use the Quarkus platform BOM (default for projects generated), add the Quarkus Langchain4J BOM and all dependency versions will align:

    <dependencyManagement>
        <dependencies>
            <dependency>
                <groupId>${quarkus.platform.group-id}</groupId>
                <artifactId>${quarkus.platform.artifact-id}</artifactId>
                <version>${quarkus.platform.version}</version>
                <type>pom</type>
                <scope>import</scope>
            </dependency>
            <dependency>
                <groupId>${quarkus.platform.group-id}</groupId>
                <artifactId>quarkus-langchain4j-bom</artifactId> (1)
                <version>${quarkus.platform.version}</version> (2)
                <type>pom</type>
                <scope>import</scope>
            </dependency>
        </dependencies>
    </dependencyManagement>

    <dependencies>
      <dependency>
        <groupId>io.quarkiverse.langchain4j</groupId>
        <artifactId>quarkus-langchain4j-jitllm</artifactId>
        (3)
      </dependency>
    </dependencies>
1 In your dependencyManagement section, add the quarkus-langchain4j-bom
2 Inherit the version from your platform version
3 Voilà, no need for version alignment anymore
If no other LLM extension is configured, AI Services will automatically use the GPU-accelerated JitLLM model!

Sample implementation for ChatModel:

@Path("chat")
public class ChatLanguageModelResource {
    private final ChatModel chatModel;
    public ChatLanguageModelResource(ChatModel chatModel) {
        this.chatModel = chatModel;
    }

    @GET
    @Path("blocking")
    public String blocking() {
        return chatModel.chat("When was the nobel prize for economics first awarded?");
    }
}

Send requests to blocking endpoint:

curl http://localhost:8080/chat/blocking

Sample implementation for StreamingChatModel:

@Path("chat")
public class ChatLanguageModelResource {
    private final StreamingChatModel streamingChatModel;
    public ChatLanguageModelResource(StreamingChatModel streamingChatModel) {
        this.streamingChatModel = streamingChatModel;
    }

    @GET
    @Path("streaming")
    @RestStreamElementType(MediaType.TEXT_PLAIN)
    public Multi<String> streaming() {
        return Multi.createFrom().emitter(emitter -> {
            streamingChatModel.chat("When was the nobel prize for economics first awarded?",
                    new StreamingChatResponseHandler() {
                        @Override
                        public void onPartialResponse(String token) {
                            emitter.emit(token);
                        }

                        @Override
                        public void onError(Throwable error) {
                            emitter.fail(error);
                        }

                        @Override
                        public void onCompleteResponse(ChatResponse completeResponse) {
                            emitter.complete();
                        }
                    });
        });
    }
}

Send requests to streaming endpoint:

curl http://localhost:8080/chat/streaming

Configure JitLLM

The JitLLM extension can be configured via standard Quarkus properties:

# Enable JitLLM integration
quarkus.langchain4j.jitllm.enable-integration=true

# Select the default model
quarkus.langchain4j.jitllm.chat-model.model-name=unsloth/Llama-3.2-1B-Instruct-GGUF
quarkus.langchain4j.jitllm.chat-model.quantization=F16
quarkus.langchain4j.jitllm.chat-model.temperature=0.7
quarkus.langchain4j.jitllm.chat-model.max-tokens=1024

Model files are automatically downloaded from Beehive Lab HuggingFace if not available locally.

To use a GGUF file you already have, point quarkus.langchain4j.jitllm.models-path at a directory laid out the way the cache is: <owner>_<name>/<name>-<quantization>.gguf and an empty .finished marker next to it. For example, local_Qwen3-4B/Qwen3-4B-Q8_0.gguf is selected with model-name=local/Qwen3-4B and quantization=Q8_0. The model name, quantization and models path are fixed at build time.

Building & Running the Quarkus Application

Dev Mode

To run your Quarkus application in dev mode with TornadoVM:

  1. Ensure your pom.xml contains the quarkus-langchain4j-jitllm dependency (shown earlier).

  2. Add the TornadoVM argfile as a Maven property:

    <properties>
        <tornado.argfile>${env.TORNADOVM_HOME}/tornado-argfile</tornado.argfile>
    </properties>
  3. Pass the argfile to the JVM in the plugin configuration for dev mode:

    <plugin>
        <groupId>io.quarkus</groupId>
        <artifactId>quarkus-maven-plugin</artifactId>
        <configuration>
            <jvmArgs>@${tornado.argfile}</jvmArgs>
        </configuration>
    </plugin>
  4. Launch dev mode explicitly:

    mvn quarkus:dev

Production Mode

To build and run your application in production mode:

  1. Build the Quarkus application:

    mvn clean package
  2. Run the generated jar with the TornadoVM argfile:

    java @$TORNADOVM_HOME/tornado-argfile --add-modules jdk.incubator.vector \
        -jar target/quarkus-app/quarkus-run.jar
Ensure TORNADOVM_HOME is set and points at the SDK matching your JDK. The tornado-argfile does not add jdk.incubator.vector, which the engine needs, so pass --add-modules jdk.incubator.vector yourself.
The model’s device-memory budget is set by quarkus.langchain4j.jitllm.chat-model.device-memory (default 4GB). The extension sets tornado.device.memory from it at model initialization, so a bare -Dtornado.device.memory on the command line has no effect.

Supported Models and Quantizations

The following models have been tested with JitLLM and can be found in Beehive Lab’s HuggingFace Collections.

Quantization format names are model-dependent. Some models require F16, others fp16 (lowercase), as defined in their GGUF files. Use exactly the spelling listed below, or the model may fail to load.

Model Quantizations Model Identifier (Hugging Face)

Llama 3.2 1B

F16, Q8_0

unsloth/Llama-3.2-1B-Instruct-GGUF

Llama 3.2 3B

F16, Q8_0

unsloth/Llama-3.2-3B-Instruct-GGUF

Llama 3.1 8B

fp16, Q8_0

brittlewis12/Meta-Llama-3.1-8B-Instruct-GGUF

Mistral 7B

fp16, Q8_0

MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF

Qwen2.5 0.5B

F16, Q8_0

bartowski/Qwen2.5-0.5B-Instruct-GGUF

Qwen2.5 1.5B

fp16, Q8_0

Qwen/Qwen2.5-1.5B-Instruct-GGUF

DeepSeek-R1 Distill Qwen 1.5B

F16, Q8_0

hdnh2006/DeepSeek-R1-Distill-Qwen-1.5B-GGUF

DeepSeek-R1 Distill Qwen 7B

F16, Q8_0

XelotX/DeepSeek-R1-Distill-Qwen-7B-GGUF

Qwen3 0.6B

F16, Q8_0

ggml-org/Qwen3-0.6B-GGUF

Qwen3 1.7B

F16, Q8_0

ggml-org/Qwen3-1.7B-GGUF

Qwen3 4B

F16, Q8_0

ggml-org/Qwen3-4B-GGUF

Qwen3 8B

F16, Q8_0

ggml-org/Qwen3-8B-GGUF

Phi 3 Mini 4k

fp16

microsoft/Phi-3-mini-4k-instruct-gguf

Phi 3 Mini 4k

Q8_0

bartowski/Phi-3-mini-4k-instruct-GGUF

Phi 3 Mini 128k

Q8_0

QuantFactory/Phi-3-mini-4k-instruct-GGUF

Phi 3.1 Mini 128k

Q8_0

bartowski/Phi-3.1-mini-128k-instruct-GGUF

Granite 3.2 2B

F16, Q8_0

ibm-research/granite-3.2-2b-instruct-GGUF

Granite 3.2 8B

F16, Q8_0

ibm-research/granite-3.2-8b-instruct-GGUF

Granite 3.3 2B

F16, Q8_0

ibm-research/granite-3.3-2b-instruct-GGUF

Granite 3.3 8B

F16, Q8_0

ibm-research/granite-3.3-8b-instruct-GGUF

Granite 4.0 1B

F16, Q8_0

ibm-research/granite-4.0-1b-instruct-GGUF

Each entry corresponds to a GGUF model tested to run on TornadoVM via jitllm.

Configuration Reference

Configuration property fixed at build time - All other configuration properties are overridable at runtime

Configuration property

Type

Default

Determines whether the necessary JitLLM models are downloaded and included in the jar at build time. Currently, this option is only valid for fast-jar deployments.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_INCLUDE_MODELS_IN_ARTIFACT

boolean

true

Whether the model should be enabled

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_ENABLED

boolean

true

Model name to use

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_MODEL_NAME

string

unsloth/Llama-3.2-1B-Instruct-GGUF

Quantization of the model to use

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_QUANTIZATION

string

F16

Location on the file-system which serves as a cache for the models

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_MODELS_PATH

path

${user.home}/.langchain4j/models

What sampling temperature to use, between 0.0 and 1.0.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_TEMPERATURE

double

0.3

What sampling topP to use, between 0.0 and 1.0.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_TOP_P

double

0.85

What seed value to use.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_SEED

int

1234

The maximum number of tokens to generate in the completion.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_MAX_TOKENS

int

512

Whether to use the prefill/decode inference path, which processes the prompt in a dedicated prefill phase before single-token decoding. Combined with a prefill-batch-size() greater than 1 this enables the batched prefill path.

Off by default. Batched prefill is default-off in the engine on every backend, and enabling it is an opt-in that needs its own performance evidence for the model and device in question.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_PREFILL_DECODE

boolean

false

Number of prompt tokens processed per batch during the prefill phase. Only used when prefill-decode() is true: a value greater than 1 enables the batched prefill path, while 1 falls back to sequential prefill/decode.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_PREFILL_BATCH_SIZE

int

1

Whether to enable the model’s thinking/reasoning phase. Only models with a thinking mode (e.g. Qwen3) honor this; for other models it has no effect.

When true, the model reasons inside <think>…</think> before answering, which tends to improve tool-calling decisions. Set to false for faster responses by skipping the reasoning phase.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_ENABLE_THINKING

boolean

false

Amount of GPU memory TornadoVM is allowed to allocate e.g. "7GB", "512MB".

Note: this maps to the JVM-global engine flag tornado.device.memory, so when multiple models are configured the first one to initialize wins.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_CHAT_MODEL_DEVICE_MEMORY

string

4GB

Whether to enable the integration. Set to false to disable all requests.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_ENABLE_INTEGRATION

boolean

true

Whether JitLLM should log requests

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_LOG_REQUESTS

boolean

false

Whether JitLLM client should log responses

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM_LOG_RESPONSES

boolean

false

Named model config

Type

Default

Model name to use

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_MODEL_NAME

string

unsloth/Llama-3.2-1B-Instruct-GGUF

Quantization of the model to use

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_QUANTIZATION

string

F16

What sampling temperature to use, between 0.0 and 1.0.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_TEMPERATURE

double

0.3

What sampling topP to use, between 0.0 and 1.0.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_TOP_P

double

0.85

What seed value to use.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_SEED

int

1234

The maximum number of tokens to generate in the completion.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_MAX_TOKENS

int

512

Whether to use the prefill/decode inference path, which processes the prompt in a dedicated prefill phase before single-token decoding. Combined with a prefill-batch-size() greater than 1 this enables the batched prefill path.

Off by default. Batched prefill is default-off in the engine on every backend, and enabling it is an opt-in that needs its own performance evidence for the model and device in question.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_PREFILL_DECODE

boolean

false

Number of prompt tokens processed per batch during the prefill phase. Only used when prefill-decode() is true: a value greater than 1 enables the batched prefill path, while 1 falls back to sequential prefill/decode.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_PREFILL_BATCH_SIZE

int

1

Whether to enable the model’s thinking/reasoning phase. Only models with a thinking mode (e.g. Qwen3) honor this; for other models it has no effect.

When true, the model reasons inside <think>…</think> before answering, which tends to improve tool-calling decisions. Set to false for faster responses by skipping the reasoning phase.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_ENABLE_THINKING

boolean

false

Amount of GPU memory TornadoVM is allowed to allocate e.g. "7GB", "512MB".

Note: this maps to the JVM-global engine flag tornado.device.memory, so when multiple models are configured the first one to initialize wins.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__CHAT_MODEL_DEVICE_MEMORY

string

4GB

Whether to enable the integration. Set to false to disable all requests.

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__ENABLE_INTEGRATION

boolean

true

Whether JitLLM should log requests

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__LOG_REQUESTS

boolean

false

Whether JitLLM client should log responses

Environment variable: QUARKUS_LANGCHAIN4J_JITLLM__MODEL_NAME__LOG_RESPONSES

boolean

false

Limitations

  • TornadoVM currently does not support GraalVM Native Image builds.

  • Ensure that the TornadoVM environment (TORNADOVM_HOME and tornado-argfile) is properly set before running Quarkus.

  • Only Java 21 or newer versions are supported.

  • Tool calling is supported for the Llama 3.x, Qwen2.5, Qwen3, Qwen3.8, Gemma 4 and Granite 3.2/4.0 families. It is not supported for Phi-3, DeepSeek-R1-Distill, Qwen MoE and Mistral/Devstral: a request that carries tools for one of these models is refused up front with an UnsupportedFeatureException, rather than sent to a model that cannot call them.

  • Forced tool choice (ToolChoice.REQUIRED) and a JSON response format are not supported; requests that set either are rejected.

  • Gemma 4 and Qwen3.8 run on both the CUDA and the OpenCL backends, as do the other families.