Skip to content

llmcompressor.transformers.finetune.data.flickr_30k

Classes:

  • Flickr30K

    :param dataset_args: configuration settings for dataset loading

Flickr30K

Flickr30K(
    dataset_args: DatasetArguments,
    split: str,
    processor: Processor,
)

Bases: TextGenerationDataset

Parameters:

  • dataset_args

    (DatasetArguments) –

    configuration settings for dataset loading

  • split

    (str) –

    split from dataset to load, for instance test or train[:5%]

  • processor

    (Processor) –

    processor or tokenizer to use on dataset

Source code in llmcompressor/transformers/finetune/data/flickr_30k.py
def __init__(
    self, dataset_args: "DatasetArguments", split: str, processor: Processor
):
    dataset_args = deepcopy(dataset_args)
    dataset_args.dataset = "lmms-lab/flickr30k"

    super().__init__(dataset_args=dataset_args, split=split, processor=processor)

    if (
        self.tokenizer is not None
        and getattr(self.tokenizer, "chat_template", None) is None
    ):
        # note that since tokenizer is a member of processor,
        # this change affects processor.apply_chat_template
        self.tokenizer.chat_template = self.DEFAULT_CHAT_TEMPLATE
        logger.warning(
            "tokenizer.chat_template is not set, using default chat template for "
            f"{self.__class__.__name__}"
        )