How to setup Diffusers for Stable Diffusion XL using a Nvidia Pascal GPU in Linux (Ubuntu or Debian based)

You need to first ensure that all the necessary drivers (CUDA most important) for the Nvidia GPU are installed. This post will NOT cover how to do that.

What i am working with here is a Nvidia GTX 1070, just so you know.

Preparing the project directory and virtual environment

Create a new directory/folder (whichever terminology your familiar with) and cd to it from a terminal instance.

Now pay attention here.

Diffusers uses PyTorch as its backend. PyTorch has several wheel repositories for different CUDA versions, and you need to choose for Pascal GPUs the one compatible with CUDA 12.1 or 12.4.

To get the right working libraries, we need to fetch them from the repository 'https://download.pytorch.org/whl/cu124' using the --index-url pip CLI parameter/argument.

Now, to the Diffusers Framework.

Once the PyTorch things is done, the only thing to watch out for is ensuring transformers is fixed at version 5.5.0 . Why is that? My hunch (and days of trying to figure out why my setup broke when i updated my libraries in my virtual environment), parts of the transformers API was changed in later version, and diffusers didn't take that into account in their newer versions. From my findings, version 5.5.0 was the last working one. Since we want this to just work, make sure that you only get THAT version until diffusers decides to fix this.

Now, lets get to what to run inside our project directory.

python3 -m venv ./venv/
./venv/bin/pip install torch torchvision --index-url 'https://download.pytorch.org/whl/cu124'
./venv/bin/pip install diffusers transformers==5.5.0 accelerate peft

This should be all that is needed to get a working virtual environment setup.

Using Diffusers for running Stable Diffusion XL models

This is now the part in which we need to interact with the Diffusers API. In this case, a regular Python script will do the job, and so, i build one for my needs.

#!./venv/bin/python3
import argparse
from pathlib import Path

import torch
from diffusers import (
    DPMSolverMultistepScheduler,
    StableDiffusionXLPipeline,
)


# Constants.
# ==========
IMAGES_BATCH_COUNT = 2
TORCH_DTYPE = torch.bfloat16


# Prepare and parse the terminal arguments.
# =========================================
args_parser = argparse.ArgumentParser(
    description="Custom program for text-to-image using Stable Diffusion XL models."
)
args_parser.add_argument(
    "--model",
    type=str,
    required=True,
    help="Huggingface repository path to the SDXL model weights.",
)
args_parser.add_argument(
    "--prompt-path",
    type=Path,
    default=Path("./prompt.txt"),
    help="Filepath to the prompt on what to generate.",
)
args_parser.add_argument(
    "--negative-prompt-path",
    type=Path,
    default=None,
    help="Filepath to the prompt on what NOT to generate. Ignored if not set.",
)
args_parser.add_argument(
    "--sampling-steps",
    type=int,
    default=26,
    help="How many steps are performed in the denoising procedure. More usually result in better image quality.",
)
args_parser.add_argument(
    "--guidance-scale",
    type=float,
    default=6.5,
    help="How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.",
)
args_parser.add_argument(
    "--images-count",
    type=int,
    default=IMAGES_BATCH_COUNT,
    help="How many images should be generated per prompt.",
)
args_parser.add_argument(
    "--seed",
    type=int,
    default=None,
    help="The seed used to set the state of the PRNG. If not set, chooses a random one.",
)
args_parser.add_argument(
    "--output-dirpath",
    type=Path,
    default=Path("./txt2img_batch"),
    help="Where to store the generated images.",
)


parsed_args = args_parser.parse_args()


# Prepare the text-to-image pipeline to generate images.
# ======================================================
prompt = parsed_args.prompt_path.read_text()
negative_prompt = None if parsed_args.negative_prompt_path == None else parsed_args.negative_prompt_path.read_text()

pipeline = StableDiffusionXLPipeline.from_single_file(
    parsed_args.model,
    custom_pipeline="lpw_stable_diffusion_xl",
    torch_dtype=TORCH_DTYPE,
    variant="bfp16",
    safety_checker=None,
)
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(
    pipeline.scheduler.config,
    use_karras_sigmas=True,
    algorithm_type="dpmsolver++",
)
pipeline.enable_model_cpu_offload()
pipeline.enable_attention_slicing(slice_size="auto")
pipeline.vae.enable_slicing()

# Generate and save the images.
# =============================
batch_count, leftover_count = divmod(parsed_args.images_count, IMAGES_BATCH_COUNT)
images_batch_sequence = []

if batch_count > 0:
    images_batch_sequence += [IMAGES_BATCH_COUNT] * batch_count
if leftover_count > 0:
    images_batch_sequence += [leftover_count]

generator = None if parsed_args.seed == None else torch.manual_seed(parsed_args.seed)

parsed_args.output_dirpath.mkdir(
    parents=True,
    exist_ok=True,
)
for num_images_per_prompt in images_batch_sequence:
    inference_result = pipeline(
        prompt=prompt,
        negative_prompt=negative_prompt,
        num_inference_steps=parsed_args.sampling_steps,
        guidance_scale=parsed_args.guidance_scale,
        num_images_per_prompt=num_images_per_prompt,
        generator=generator,
    )

    file_counter = 0
    saved_image_paths = []
    for generated_image in inference_result.images:
        output_filepath = parsed_args.output_dirpath.joinpath(f"image{file_counter}.png")
        while output_filepath.exists():
            file_counter += 1
            output_filepath = parsed_args.output_dirpath.joinpath(f"image{file_counter}.png")

        generated_image.save(output_filepath)
        saved_image_paths.append(output_filepath.name)

    print(f"Saved the generated {num_images_per_prompt} image/s {saved_image_paths} into the directory '{parsed_args.output_dirpath}'.")

This is what i use to generate AI images with Stable Diffusion XL models. It has a solid CLI interface, and the script can be adapted for ones individual needs if something is insufficient.

One important thing to note, due to the Nvidia GTX 1070 being limited to 8 GB of VRAM, and most Stable Diffusion XL models being around 7 GB, reducing the VRAM usage is a must, and this is done usually with offloading, attention slicing and VAE slicing.

The best case for offloading is using model offloading, but you need to ensure that almost nothing is filling the GPUs VRAM, otherwise the script will likely exit with a out of memory exception when saving the generated images. If that is happening to you way to often, and you have closed all programs, change model offloading to cpu offloading. It will take longer, but crashes should happen rarely. Even better (speed and reliability wise) is using group offloading with CUDA streaming enabled, but this required having more RAM (16 GB is not enough, 24+ GB is recommended).

Regarding attention slicing, it kinda works on the Nvidia GTX 1070, but is very important to enable, since it lowers the VRAM use further, with minimal speed loss. Now, why it is kinda working? Due to FlashAttention (the enabled default on newer PyTorch builds) simply not working with my GPU, and a altenative being urgently needed. My only option (after hours of research and testing) ended up being the enable_attention_slicing method from the base class DiffusionPipeline class, which does (almost) the same thing luckily, so we are not totally fucked yet.

VAE slicing saves memory by splitting large batches of inputs into a single batch of data and separately processes them. This method works best when generating more than one image at a time, which will be the case most of the time, but keep in mind that inference will take longer the more images you generate.

Generating images faster using Ays

Now this is mostly uncharted territory, but i was researching how to speed up the inference of Stable Diffusion XL models on my GPU, and found this gem.

Now, how useful is using the AYS schedule for inference? In my experience, it is so-so.

Yes, its much faster and does generate good enough images. But...the image quality in regards to prompt adherence is worse, and this depends alot on the model your using. If it was fine tuned very well, it will perform mostly fine. If not, the images will most likely look weird.

Also (from my findings), it is best to use the sde-dpmsolver++ scheduler, due to it converging faster and giving good quality results, in 10 steps of inference.

Probably best used for experimenting with new ideas.

Here is the (mostly) same script, but using AYS for the inference.

#!./venv/bin/python3
import argparse
from pathlib import Path

import torch
from diffusers import (
    DPMSolverMultistepScheduler,
    StableDiffusionXLPipeline,
)
from diffusers.schedulers import AysSchedules


# Constants.
# ==========
IMAGES_BATCH_COUNT = 2
TORCH_DTYPE = torch.bfloat16


# Prepare and parse the terminal arguments.
# =========================================
args_parser = argparse.ArgumentParser(
    description="Custom program for text-to-image using normal Stable Diffusion models."
)
args_parser.add_argument(
    "--model",
    type=str,
    required=True,
    help="Huggingface repository path to the SDXL model weights.",
)
args_parser.add_argument(
    "--prompt-path",
    type=Path,
    default=Path("./prompt.txt"),
    help="Filepath to the prompt on what to generate.",
)
args_parser.add_argument(
    "--negative-prompt-path",
    type=Path,
    default=None,
    help="Filepath to the prompt on what NOT to generate. Ignored if not set.",
)
args_parser.add_argument(
    "--guidance-scale",
    type=float,
    default=3.0,
    help="How close it should follow the prompt. Lower values result in more freedom, while higher ones result in more consistency.",
)
args_parser.add_argument(
    "--images-count",
    type=int,
    default=IMAGES_BATCH_COUNT,
    help="How many images should be generated per prompt.",
)
args_parser.add_argument(
    "--seed",
    type=int,
    default=None,
    help="The seed used to set the state of the PRNG. If not set, chooses a random one.",
)
args_parser.add_argument(
    "--output-dirpath",
    type=Path,
    default=Path("./txt2img_batch"),
    help="Where to store the generated images.",
)


parsed_args = args_parser.parse_args()


# Prepare the text-to-image pipeline to generate images.
# ======================================================
prompt = parsed_args.prompt_path.read_text()
negative_prompt = None if parsed_args.negative_prompt_path == None else parsed_args.negative_prompt_path.read_text()

pipeline = StableDiffusionXLPipeline.from_single_file(
    parsed_args.model,
    custom_pipeline="lpw_stable_diffusion_xl",
    torch_dtype=TORCH_DTYPE,
    safety_checker=None,
)
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(
    pipeline.scheduler.config,
    algorithm_type="sde-dpmsolver++",
    solver_order=2,
)
pipeline.enable_model_cpu_offload()
pipeline.enable_attention_slicing(slice_size="auto")
pipeline.vae.enable_slicing()


# Generate and save the images.
# =============================
batch_count, leftover_count = divmod(parsed_args.images_count, IMAGES_BATCH_COUNT)
images_batch_sequence = []

if batch_count > 0:
    images_batch_sequence += [IMAGES_BATCH_COUNT] * batch_count
if leftover_count > 0:
    images_batch_sequence += [leftover_count]

generator = None if parsed_args.seed == None else torch.manual_seed(parsed_args.seed)

parsed_args.output_dirpath.mkdir(
    parents=True,
    exist_ok=True,
)
for num_images_per_prompt in images_batch_sequence:
    inference_result = pipeline(
        prompt=prompt,
        negative_prompt=negative_prompt,
        guidance_scale=parsed_args.guidance_scale,
        num_images_per_prompt=num_images_per_prompt,
        generator=generator,
        timesteps=AysSchedules["StableDiffusionXLTimesteps"],
    )

    file_counter = 0
    for generated_image in inference_result.images:
        output_filepath = parsed_args.output_dirpath.joinpath(f"image{file_counter}.png")
        while output_filepath.exists():
            file_counter += 1
            output_filepath = parsed_args.output_dirpath.joinpath(f"image{file_counter}.png")

        generated_image.save(output_filepath)

    print(f"Saved the generated {num_images_per_prompt} image/s into the directory '{parsed_args.output_dirpath}'.")

Conclusion

It's doable, but will get harder with time, due to lacking software support, and whatever currently works will likely no longer be available in the near future, unless you build the libraries yourself (and even that is sometimes a PITA). How do i know that? Why am i this opinion? cuML and cuDF. Cannot get binary wheels that have Pascal support, and building them locally failed with cryptic error messages (maybe my setup sucks, but it is what it is at the moment). Best i got running is cupy for GPGPU things.

So, if you wanna do it, you still can, but do not expect miracles, and inference on the Nvidia 1000 series is not the best (the GTX 1070 needs around 3:30 minutes for generating 2 images with model offloading enabled, to give you a idea).

links

social