콘텐츠로 이동

Gemma4

Gemma4는 텍스트와 이미지를 입력으로 처리하고 텍스트를 출력으로 생성하는 멀티모달 모델로, 사전 학습된 변형과 지침 조정된 변형 모두에 대해 오픈 웨이트를 제공합니다. Dense 및 MoE 아키텍처를 모두 갖춘 Gemma4는 텍스트 생성, 코딩, 추론과 같은 작업에 적합합니다. RBLN NPU는 Optimum RBLN을 사용하여 Gemma4 모델 추론을 가속화할 수 있습니다.

API Reference

Classes

RBLNGemma4VisionModel

Bases: RBLNModel

Gemma4 vision encoder model optimized for RBLN NPU.

This model inherits from [RBLNModel]. It implements the methods to convert and run pre-trained transformers based Gemma4 vision encoder model on RBLN devices by:

  • transferring the checkpoint weights of the original into an optimized RBLN graph,
  • compiling the resulting graph using the RBLN compiler.

patch_embedder (per-patch linear projection + 2D position embedding lookup) and rotary_emb (multidimensional cos/sin tables) both run on the host (CPU). patch_embedder weights are persisted as a saved torch artifact; rotary_emb is recreated from config since its inv_freq buffer is non-persistent. The compiled Gemma4VisionModelWrapper (encoder-layers -> pooler) takes the host-computed inputs_embeds, pixel_position_ids, and (cos, sin) rotary tables as inputs. Padding within max_patches is handled by the encoder via pixel_position_ids == -1 markers.

Methods:

from_model(model, config=None, rbln_config=None, model_save_dir=None, subfolder='', **kwargs) classmethod

Converts and compiles a pre-trained HuggingFace library model into a RBLN model. This method performs the actual model conversion and compilation process.

Parameters:

Name Type Description Default
model PreTrainedModel

The PyTorch model to be compiled. The object must be an instance of the HuggingFace transformers PreTrainedModel class.

required
config PretrainedConfig | None

The configuration object associated with the model.

None
rbln_config RBLNModelConfig | dict | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

The method performs the following steps:

  1. Compiles the PyTorch model into an optimized RBLN graph
  2. Configures the model for the specified NPU device
  3. Creates the necessary runtime objects if requested
  4. Saves the compiled model and configurations

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

from_pretrained(model_id, export=None, rbln_config=None, **kwargs) classmethod

The from_pretrained() function is utilized in its standard form as in the HuggingFace transformers library. User can use this function to load a pre-trained model from the HuggingFace library and convert it to a RBLN model to be run on RBLN NPUs.

Parameters:

Name Type Description Default
model_id str | Path

The model id of the pre-trained model to be loaded. It can be downloaded from the HuggingFace model hub or a local path, or a model id of a compiled model using the RBLN Compiler.

required
export bool | None

A boolean flag to indicate whether the model should be compiled. If None, it will be determined based on the existence of the compiled model files in the model_id.

None
rbln_config dict | RBLNModelConfig | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

save_pretrained(save_directory, push_to_hub=False, **kwargs)

Saves a model and its configuration file to a directory, so that it can be re-loaded using the [~optimum.rbln.modeling_base.RBLNBaseModel.from_pretrained] class method.

Parameters:

Name Type Description Default
save_directory str | Path

Directory where to save the model file.

required
push_to_hub bool

Whether or not to push your model to the HuggingFace model hub after saving it.

False

RBLNGemma4ForCausalLM

Bases: RBLNDecoderOnlyModelForCausalLM

Gemma4 model with a causal language modeling head optimized for RBLN NPU.

This model inherits from [RBLNModel]. It implements the methods to convert and run pre-trained transformers based Gemma4ForCausalLM model on RBLN devices by:

  • transferring the checkpoint weights of the original into an optimized RBLN graph,
  • compiling the resulting graph using the RBLN compiler.

Compared to the base decoder-only class, this class additionally saves and loads embed_tokens_per_layer (the auxiliary per-layer-input embedding) alongside embed_tokens as a torch artifact, and wires the resulting per_layer_inputs tensor through the runtime to match the Gemma4ForCausalLMWrapper argument order.

Methods:

generate(input_ids, attention_mask=None, generation_config=None, **kwargs)

The generate function is utilized in its standard form as in the HuggingFace transformers library. User can use this function to generate text from the model. Check the HuggingFace transformers documentation for more details.

Parameters:

Name Type Description Default
input_ids LongTensor

The input ids to the model.

required
attention_mask LongTensor

The attention mask to the model.

None
generation_config GenerationConfig

The generation configuration to be used as base parametrization for the generation call. **kwargs passed to generate matching the attributes of generation_config will override them. If generation_config is not provided, the default will be used, which had the following loading priority: 1) from the generation_config.json model file, if it exists; 2) from the model configuration. Please note that unspecified parameters will inherit GenerationConfig’s default values.

None
kwargs dict[str, Any]

Additional arguments passed to the generate function. See the HuggingFace transformers documentation for more details.

{}

Returns:

Type Description
ModelOutput | LongTensor

A ModelOutput (if return_dict_in_generate=True or when config.return_dict_in_generate=True) or a torch.LongTensor.

from_pretrained(model_id, export=None, rbln_config=None, **kwargs) classmethod

The from_pretrained() function is utilized in its standard form as in the HuggingFace transformers library. User can use this function to load a pre-trained model from the HuggingFace library and convert it to a RBLN model to be run on RBLN NPUs.

Parameters:

Name Type Description Default
model_id str | Path

The model id of the pre-trained model to be loaded. It can be downloaded from the HuggingFace model hub or a local path, or a model id of a compiled model using the RBLN Compiler.

required
export bool | None

A boolean flag to indicate whether the model should be compiled. If None, it will be determined based on the existence of the compiled model files in the model_id.

None
rbln_config dict | RBLNModelConfig | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

save_pretrained(save_directory, push_to_hub=False, **kwargs)

Saves a model and its configuration file to a directory, so that it can be re-loaded using the [~optimum.rbln.modeling_base.RBLNBaseModel.from_pretrained] class method.

Parameters:

Name Type Description Default
save_directory str | Path

Directory where to save the model file.

required
push_to_hub bool

Whether or not to push your model to the HuggingFace model hub after saving it.

False
from_model(model, config=None, rbln_config=None, model_save_dir=None, subfolder='', **kwargs) classmethod

Converts and compiles a pre-trained HuggingFace library model into a RBLN model. This method performs the actual model conversion and compilation process.

Parameters:

Name Type Description Default
model PreTrainedModel

The PyTorch model to be compiled. The object must be an instance of the HuggingFace transformers PreTrainedModel class.

required
config PretrainedConfig | None

The configuration object associated with the model.

None
rbln_config RBLNModelConfig | dict | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

The method performs the following steps:

  1. Compiles the PyTorch model into an optimized RBLN graph
  2. Configures the model for the specified NPU device
  3. Creates the necessary runtime objects if requested
  4. Saves the compiled model and configurations

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

set_adapter(adapter_name)

Sets the active adapter(s) for the model using adapter name(s).

Parameters:

Name Type Description Default
adapter_name str | list[str]

The name(s) of the adapter(s) to be activated. Can be a single adapter name or a list of adapter names.

required

Raises:

Type Description
ValueError

If the model is not configured with LoRA or if the adapter name is not found.

RBLNGemma4ForConditionalGeneration

Bases: RBLNModel, RBLNDecoderOnlyGenerationMixin

Gemma4 model for image-text-to-text generation optimized for RBLN NPU.

This model inherits from [RBLNModel]. It implements the methods to convert and run pre-trained transformers based Gemma4ForConditionalGeneration model on RBLN devices by:

  • transferring the checkpoint weights of the original into an optimized RBLN graph,
  • compiling the resulting graph using the RBLN compiler.

This class compiles the embed_vision multimodal projector (vision soft tokens -> language-model embedding space) as its own graph. Vision encoding and language modeling are compiled as submodules: vision_tower ([RBLNGemma4VisionModel], batch_size=1, looped over images/frames at runtime) and language_model ([RBLNGemma4ForCausalLM]).

Both image and video inputs are supported. A video is (num_videos, num_frames, ...); get_video_features flattens the leading dims and reuses the per-image vision path, so each frame is encoded independently by the looped batch-1 vision tower. Image (token_type 1) and video (token_type 2) soft tokens are both dispatched to the bidirectional image_prefill graph (one chunk per contiguous same-token_type run), giving per-frame bidirectional attention with cross-frame causal attention. This assumes the HF token layout separates consecutive frames with text (timestamps + BOI/EOI), so each frame is its own run that fits one soft-token bucket; adjacent frames with no separating text would merge into a single over-long run. Audio inputs are not supported.

Methods:

from_model(model, config=None, rbln_config=None, model_save_dir=None, subfolder='', **kwargs) classmethod

Converts and compiles a pre-trained HuggingFace library model into a RBLN model. This method performs the actual model conversion and compilation process.

Parameters:

Name Type Description Default
model PreTrainedModel

The PyTorch model to be compiled. The object must be an instance of the HuggingFace transformers PreTrainedModel class.

required
config PretrainedConfig | None

The configuration object associated with the model.

None
rbln_config RBLNModelConfig | dict | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

The method performs the following steps:

  1. Compiles the PyTorch model into an optimized RBLN graph
  2. Configures the model for the specified NPU device
  3. Creates the necessary runtime objects if requested
  4. Saves the compiled model and configurations

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

generate(input_ids, attention_mask=None, generation_config=None, **kwargs)

The generate function is utilized in its standard form as in the HuggingFace transformers library. User can use this function to generate text from the model. Check the HuggingFace transformers documentation for more details.

Parameters:

Name Type Description Default
input_ids LongTensor

The input ids to the model.

required
attention_mask LongTensor

The attention mask to the model.

None
generation_config GenerationConfig

The generation configuration to be used as base parametrization for the generation call. **kwargs passed to generate matching the attributes of generation_config will override them. If generation_config is not provided, the default will be used, which had the following loading priority: 1) from the generation_config.json model file, if it exists; 2) from the model configuration. Please note that unspecified parameters will inherit GenerationConfig’s default values.

None
kwargs dict[str, Any]

Additional arguments passed to the generate function. See the HuggingFace transformers documentation for more details.

{}

Returns:

Type Description
ModelOutput | LongTensor

A ModelOutput (if return_dict_in_generate=True or when config.return_dict_in_generate=True) or a torch.LongTensor.

from_pretrained(model_id, export=None, rbln_config=None, **kwargs) classmethod

The from_pretrained() function is utilized in its standard form as in the HuggingFace transformers library. User can use this function to load a pre-trained model from the HuggingFace library and convert it to a RBLN model to be run on RBLN NPUs.

Parameters:

Name Type Description Default
model_id str | Path

The model id of the pre-trained model to be loaded. It can be downloaded from the HuggingFace model hub or a local path, or a model id of a compiled model using the RBLN Compiler.

required
export bool | None

A boolean flag to indicate whether the model should be compiled. If None, it will be determined based on the existence of the compiled model files in the model_id.

None
rbln_config dict | RBLNModelConfig | None

Configuration for RBLN model compilation and runtime. This can be provided as a dictionary or an instance of the model's configuration class (e.g., RBLNLlamaForCausalLMConfig for Llama models). For detailed configuration options, see the specific model's configuration class documentation.

None
kwargs Any

Additional keyword arguments. Arguments with the prefix rbln_ are passed to rbln_config, while the remaining arguments are passed to the HuggingFace library.

{}

Returns:

Type Description
RBLNModel

A RBLN model instance ready for inference on RBLN NPU devices.

save_pretrained(save_directory, push_to_hub=False, **kwargs)

Saves a model and its configuration file to a directory, so that it can be re-loaded using the [~optimum.rbln.modeling_base.RBLNBaseModel.from_pretrained] class method.

Parameters:

Name Type Description Default
save_directory str | Path

Directory where to save the model file.

required
push_to_hub bool

Whether or not to push your model to the HuggingFace model hub after saving it.

False

Functions:

Classes

RBLNGemma4ForCausalLMConfig

Bases: RBLNDecoderOnlyModelForCausalLMConfig

Attributes

is_auto_num_blocks property

Returns True if kvcache_num_blocks will be automatically determined during compilation to fit within the available DRAM on the NPU.

Methods:

__init__(use_position_ids=None, use_attention_mask=None, prefill_chunk_size=None, image_prefill_chunk_size=None, **kwargs)

Parameters:

Name Type Description Default
use_position_ids bool | None

Whether to use position_ids. Forced to True for Gemma4.

None
use_attention_mask bool | None

Whether to use attention_mask. Forced to True for Gemma4.

None
prefill_chunk_size int | None

Chunk size used during the prefill phase. When unset, it is resolved at compile time to 512 on RBLN-CR NPUs and 128 otherwise.

None
image_prefill_chunk_size int | list[int] | None

Chunk size(s) used for image-prefill

(multimodal Gemma4). A single int compiles one image_prefill graph; a list compiles one graph per value (sorted in descending order) and the runtime picks the smallest bucket that fits an image run. When not given, it is derived from the vision tower's max_soft_tokens (the default single max_soft_tokens=280 yields a single chunk size of 384).

None
kwargs Any

Additional arguments passed to the parent RBLNDecoderOnlyModelForCausalLMConfig.

{}

Raises:

Type Description
ValueError

If use_attention_mask or use_position_ids are False, or if any image-prefill chunk size is not a positive integer divisible by 128.

from_pretrained(path, rbln_config=None, return_unused_kwargs=False, **kwargs) classmethod

Load a RBLNModelConfig from a path.

Parameters:

Name Type Description Default
path str

Path to the RBLNModelConfig file or directory containing the config file.

required
rbln_config dict[str, Any] | None

Additional configuration to override.

None
return_unused_kwargs bool

Whether to return unused kwargs.

False
kwargs dict[str, Any] | None

Additional keyword arguments to override configuration values. Keys starting with 'rbln_' will have the prefix removed and be used to update the configuration.

{}

Returns:

Name Type Description
RBLNModelConfig Union[RBLNModelConfig, tuple[RBLNModelConfig, dict[str, Any]]]

The loaded configuration instance.

Note

This method loads the configuration from the specified path and applies any provided overrides. If the loaded configuration class doesn't match the expected class, a warning will be logged.

Examples:

config = RBLNResNetForImageClassificationConfig.from_pretrained("/path/to/model")

RBLNGemma4VisionModelConfig

Bases: RBLNModelConfig

Methods:

__init__(batch_size=None, max_soft_tokens=None, pooling_kernel_size=None, patch_size=None, output_hidden_states=None, **kwargs)

Parameters:

Name Type Description Default
batch_size int | None

The batch size of images (number of images, not patches). Defaults to 1.

None
max_soft_tokens int | list[int] | None

The number of soft tokens emitted per image after pooling. Defaults to 280 (the upstream default in Gemma4ImageProcessor). A single int compiles one vision graph; a list compiles one graph per value (sorted descending) so the runtime can serve images at multiple soft-token counts. Must be a value supported by the image processor (e.g. 70/140/280/560/1120).

None
pooling_kernel_size int | None

Spatial pooling kernel size applied after patchification. Defaults to model_config.pooling_kernel_size (3 by default).

None
patch_size int | None

Patch height/width in pixels. Defaults to model_config.patch_size.

None
output_hidden_states bool | None

Whether to return per-layer hidden states.

None
kwargs Any

Additional arguments passed to the parent RBLNModelConfig.

{}

Raises:

Type Description
ValueError

If batch_size is not a positive integer.

from_pretrained(path, rbln_config=None, return_unused_kwargs=False, **kwargs) classmethod

Load a RBLNModelConfig from a path.

Parameters:

Name Type Description Default
path str

Path to the RBLNModelConfig file or directory containing the config file.

required
rbln_config dict[str, Any] | None

Additional configuration to override.

None
return_unused_kwargs bool

Whether to return unused kwargs.

False
kwargs dict[str, Any] | None

Additional keyword arguments to override configuration values. Keys starting with 'rbln_' will have the prefix removed and be used to update the configuration.

{}

Returns:

Name Type Description
RBLNModelConfig Union[RBLNModelConfig, tuple[RBLNModelConfig, dict[str, Any]]]

The loaded configuration instance.

Note

This method loads the configuration from the specified path and applies any provided overrides. If the loaded configuration class doesn't match the expected class, a warning will be logged.

Examples:

config = RBLNResNetForImageClassificationConfig.from_pretrained("/path/to/model")

RBLNGemma4ForConditionalGenerationConfig

Bases: RBLNModelConfig

Methods:

__init__(batch_size=None, vision_tower=None, language_model=None, **kwargs)

Parameters:

Name Type Description Default
batch_size int | None

The batch size for inference. Defaults to 1.

None
vision_tower RBLNModelConfig | None

Configuration for the vision encoder component.

None
language_model RBLNModelConfig | None

Configuration for the language model component.

None
kwargs Any

Additional arguments passed to the parent RBLNModelConfig.

{}

Raises:

Type Description
ValueError

If batch_size is not a positive integer.

from_pretrained(path, rbln_config=None, return_unused_kwargs=False, **kwargs) classmethod

Load a RBLNModelConfig from a path.

Parameters:

Name Type Description Default
path str

Path to the RBLNModelConfig file or directory containing the config file.

required
rbln_config dict[str, Any] | None

Additional configuration to override.

None
return_unused_kwargs bool

Whether to return unused kwargs.

False
kwargs dict[str, Any] | None

Additional keyword arguments to override configuration values. Keys starting with 'rbln_' will have the prefix removed and be used to update the configuration.

{}

Returns:

Name Type Description
RBLNModelConfig Union[RBLNModelConfig, tuple[RBLNModelConfig, dict[str, Any]]]

The loaded configuration instance.

Note

This method loads the configuration from the specified path and applies any provided overrides. If the loaded configuration class doesn't match the expected class, a warning will be logged.

Examples:

config = RBLNResNetForImageClassificationConfig.from_pretrained("/path/to/model")

Functions: