diffusers/scripts/convert_original_stable_dif...

# coding=utf-8
# Copyright 2022 The HuggingFace Inc. team.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#     http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
""" Conversion script for the LDM checkpoints. """

import argparse

from diffusers.pipelines.stable_diffusion.convert_from_ckpt import load_pipeline_from_original_stable_diffusion_ckpt


if __name__ == "__main__":
    parser = argparse.ArgumentParser()

    parser.add_argument(
        "--checkpoint_path", default=None, type=str, required=True, help="Path to the checkpoint to convert."
    )
    # !wget https://raw.githubusercontent.com/CompVis/stable-diffusion/main/configs/stable-diffusion/v1-inference.yaml
    parser.add_argument(
        "--original_config_file",
        default=None,
        type=str,
        help="The YAML config file corresponding to the original architecture.",
    )
    parser.add_argument(
        "--num_in_channels",
        default=None,
        type=int,
        help="The number of input channels. If `None` number of input channels will be automatically inferred.",
    )
    parser.add_argument(
        "--scheduler_type",
        default="pndm",
        type=str,
        help="Type of scheduler to use. Should be one of ['pndm', 'lms', 'ddim', 'euler', 'euler-ancestral', 'dpm']",
    )
    parser.add_argument(
        "--pipeline_type",
        default=None,
        type=str,
        help=(
            "The pipeline type. One of 'FrozenOpenCLIPEmbedder', 'FrozenCLIPEmbedder', 'PaintByExample'"
            ". If `None` pipeline will be automatically inferred."
        ),
    )
    parser.add_argument(
        "--image_size",
        default=None,
        type=int,
        help=(
            "The image size that the model was trained on. Use 512 for Stable Diffusion v1.X and Stable Siffusion v2"
            " Base. Use 768 for Stable Diffusion v2."
        ),
    )
    parser.add_argument(
        "--prediction_type",
        default=None,
        type=str,
        help=(
            "The prediction type that the model was trained on. Use 'epsilon' for Stable Diffusion v1.X and Stable"
            " Diffusion v2 Base. Use 'v_prediction' for Stable Diffusion v2."
        ),
    )
    parser.add_argument(
        "--extract_ema",
        action="store_true",
        help=(
            "Only relevant for checkpoints that have both EMA and non-EMA weights. Whether to extract the EMA weights"
            " or not. Defaults to `False`. Add `--extract_ema` to extract the EMA weights. EMA weights usually yield"
            " higher quality images for inference. Non-EMA weights are usually better to continue fine-tuning."
        ),
    )
    parser.add_argument(
        "--upcast_attention",
        action="store_true",
        help=(
            "Whether the attention computation should always be upcasted. This is necessary when running stable"
            " diffusion 2.1."
        ),
    )
    parser.add_argument(
        "--from_safetensors",
        action="store_true",
        help="If `--checkpoint_path` is in `safetensors` format, load checkpoint with safetensors instead of PyTorch.",
    )
    parser.add_argument(
        "--to_safetensors",
        action="store_true",
        help="Whether to store pipeline in safetensors format or not.",
    )
    parser.add_argument("--dump_path", default=None, type=str, required=True, help="Path to the output model.")
    parser.add_argument("--device", type=str, help="Device to use (e.g. cpu, cuda:0, cuda:1, etc.)")
    parser.add_argument(
        "--stable_unclip",
        type=str,
        default=None,
        required=False,
        help="Set if this is a stable unCLIP model. One of 'txt2img' or 'img2img'.",
    )
    parser.add_argument(
        "--stable_unclip_prior",
        type=str,
        default=None,
        required=False,
        help="Set if this is a stable unCLIP txt2img model. Selects which prior to use. If `--stable_unclip` is set to `txt2img`, the karlo prior (https://huggingface.co/kakaobrain/karlo-v1-alpha/tree/main/prior) is selected by default.",
    )
    parser.add_argument(
        "--clip_stats_path",
        type=str,
        help="Path to the clip stats file. Only required if the stable unclip model's config specifies `model.params.noise_aug_config.params.clip_stats_path`.",
        required=False,
    )
    args = parser.parse_args()

    pipe = load_pipeline_from_original_stable_diffusion_ckpt(
        checkpoint_path=args.checkpoint_path,
        original_config_file=args.original_config_file,
        image_size=args.image_size,
        prediction_type=args.prediction_type,
        model_type=args.pipeline_type,
        extract_ema=args.extract_ema,
        scheduler_type=args.scheduler_type,
        num_in_channels=args.num_in_channels,
        upcast_attention=args.upcast_attention,
        from_safetensors=args.from_safetensors,
        device=args.device,
        stable_unclip=args.stable_unclip,
        stable_unclip_prior=args.stable_unclip_prior,
        clip_stats_path=args.clip_stats_path,
    )
    pipe.save_pretrained(args.dump_path, safe_serialization=args.to_safetensors)
Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`# coding=utf-8`
			`# Copyright 2022 The HuggingFace Inc. team.`
			`#`
			`# Licensed under the Apache License, Version 2.0 (the "License");`
			`# you may not use this file except in compliance with the License.`
			`# You may obtain a copy of the License at`
			`#`
			`# http://www.apache.org/licenses/LICENSE-2.0`
			`#`
			`# Unless required by applicable law or agreed to in writing, software`
			`# distributed under the License is distributed on an "AS IS" BASIS,`
			`# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.`
			`# See the License for the specific language governing permissions and`
			`# limitations under the License.`
			`""" Conversion script for the LDM checkpoints. """`

			`import argparse`

Module-ise "original stable diffusion to diffusers" conversion script (#2019) * convert __main__ to a function call and call it * add missing type hint * make style check pass * move loading to src/diffusers Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-01-20 09:30:44 -07:00			`from diffusers.pipelines.stable_diffusion.convert_from_ckpt import load_pipeline_from_original_stable_diffusion_ckpt`
Update conversion script to correctly handle SD 2 (#1511) * Conversion SD 2 * finish 2022-12-02 04:28:01 -07:00

Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`if __name__ == "__main__":`
			`parser = argparse.ArgumentParser()`

			`parser.add_argument(`
			`"--checkpoint_path", default=None, type=str, required=True, help="Path to the checkpoint to convert."`
			`)`
			`# !wget https://raw.githubusercontent.com/CompVis/stable-diffusion/main/configs/stable-diffusion/v1-inference.yaml`
			`parser.add_argument(`
			`"--original_config_file",`
			`default=None,`
			`type=str,`
			`help="The YAML config file corresponding to the original architecture.",`
			`)`
Add paint by example (#1533) * add paint by example * mkae loading possibel * up * Update src/diffusers/models/attention.py * up * finalize weight structure * make example work * make it work * up * up * fix * del * add * update * Apply suggestions from code review * correct transformer 2d * finish * up * up * up * up * fix * Apply suggestions from code review Co-authored-by: Pedro Cuenca <pedro@huggingface.co> * Apply suggestions from code review * up * finish Co-authored-by: Pedro Cuenca <pedro@huggingface.co> 2022-12-07 03:06:30 -07:00			`parser.add_argument(`
			`"--num_in_channels",`
			`default=None,`
			`type=int,`
			help="The number of input channels. If `None` number of input channels will be automatically inferred.",
			`)`
Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`parser.add_argument(`
			`"--scheduler_type",`
			`default="pndm",`
			`type=str,`
Correct help text for scheduler_type flag in scripts. (#1749) 2022-12-19 03:27:23 -07:00			`help="Type of scheduler to use. Should be one of ['pndm', 'lms', 'ddim', 'euler', 'euler-ancestral', 'dpm']",`
Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`)`
Add paint by example (#1533) * add paint by example * mkae loading possibel * up * Update src/diffusers/models/attention.py * up * finalize weight structure * make example work * make it work * up * up * fix * del * add * update * Apply suggestions from code review * correct transformer 2d * finish * up * up * up * up * fix * Apply suggestions from code review Co-authored-by: Pedro Cuenca <pedro@huggingface.co> * Apply suggestions from code review * up * finish Co-authored-by: Pedro Cuenca <pedro@huggingface.co> 2022-12-07 03:06:30 -07:00			`parser.add_argument(`
			`"--pipeline_type",`
			`default=None,`
			`type=str,`
misc fixes (#2282) Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-08 10:02:42 -07:00			`help=(`
			`"The pipeline type. One of 'FrozenOpenCLIPEmbedder', 'FrozenCLIPEmbedder', 'PaintByExample'"`
			". If `None` pipeline will be automatically inferred."
			`),`
Add paint by example (#1533) * add paint by example * mkae loading possibel * up * Update src/diffusers/models/attention.py * up * finalize weight structure * make example work * make it work * up * up * fix * del * add * update * Apply suggestions from code review * correct transformer 2d * finish * up * up * up * up * fix * Apply suggestions from code review Co-authored-by: Pedro Cuenca <pedro@huggingface.co> * Apply suggestions from code review * up * finish Co-authored-by: Pedro Cuenca <pedro@huggingface.co> 2022-12-07 03:06:30 -07:00			`)`
Add an explicit `--image_size` to the conversion script (#1509) * Add an explicit `--image_size` to the conversion script * style 2022-12-01 11:22:48 -07:00			`parser.add_argument(`
			`"--image_size",`
Update conversion script to correctly handle SD 2 (#1511) * Conversion SD 2 * finish 2022-12-02 04:28:01 -07:00			`default=None,`
Add an explicit `--image_size` to the conversion script (#1509) * Add an explicit `--image_size` to the conversion script * style 2022-12-01 11:22:48 -07:00			`type=int,`
			`help=(`
			`"The image size that the model was trained on. Use 512 for Stable Diffusion v1.X and Stable Siffusion v2"`
			`" Base. Use 768 for Stable Diffusion v2."`
			`),`
			`)`
Update conversion script to correctly handle SD 2 (#1511) * Conversion SD 2 * finish 2022-12-02 04:28:01 -07:00			`parser.add_argument(`
			`"--prediction_type",`
			`default=None,`
Correct type from int to str in conversion script sd 2022-12-05 11:51:29 -07:00			`type=str,`
Update conversion script to correctly handle SD 2 (#1511) * Conversion SD 2 * finish 2022-12-02 04:28:01 -07:00			`help=(`
			`"The prediction type that the model was trained on. Use 'epsilon' for Stable Diffusion v1.X and Stable"`
misc fixes (#2282) Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-08 10:02:42 -07:00			`" Diffusion v2 Base. Use 'v_prediction' for Stable Diffusion v2."`
Update conversion script to correctly handle SD 2 (#1511) * Conversion SD 2 * finish 2022-12-02 04:28:01 -07:00			`),`
			`)`
CompVis -> diffusers script - allow converting from merged checkpoint to either EMA or non-EMA (#991) * improve script * up 2022-10-26 04:32:07 -06:00			`parser.add_argument(`
			`"--extract_ema",`
			`action="store_true",`
			`help=(`
			`"Only relevant for checkpoints that have both EMA and non-EMA weights. Whether to extract the EMA weights"`
			" or not. Defaults to `False`. Add `--extract_ema` to extract the EMA weights. EMA weights usually yield"
			`" higher quality images for inference. Non-EMA weights are usually better to continue fine-tuning."`
			`),`
			`)`
Add text encoder conversion (#1559) * Initial code for attempt at improving SD <--> diffusers conversions for v2.0 * Updates to support round-trip between orig. SD 2.0 and diffusers models * Corrected formatting to Black standard * Correcting import formatting * Fixed imports (properly this time) * add some corrections * remove inference files Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-12-12 02:07:42 -07:00			`parser.add_argument(`
Fix unused upcast_attn flag in convert_original_stable_diffusion_to_diffusers script (#1942) Fix unused upcast_attn flag in sd to diffusers script 2023-01-12 11:55:40 -07:00			`"--upcast_attention",`
misc fixes (#2282) Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-08 10:02:42 -07:00			`action="store_true",`
Add text encoder conversion (#1559) * Initial code for attempt at improving SD <--> diffusers conversions for v2.0 * Updates to support round-trip between orig. SD 2.0 and diffusers models * Corrected formatting to Black standard * Correcting import formatting * Fixed imports (properly this time) * add some corrections * remove inference files Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-12-12 02:07:42 -07:00			`help=(`
			`"Whether the attention computation should always be upcasted. This is necessary when running stable"`
			`" diffusion 2.1."`
			`),`
			`)`
Module-ise "original stable diffusion to diffusers" conversion script (#2019) * convert __main__ to a function call and call it * add missing type hint * make style check pass * move loading to src/diffusers Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-01-20 09:30:44 -07:00			`parser.add_argument(`
			`"--from_safetensors",`
			`action="store_true",`
			help="If `--checkpoint_path` is in `safetensors` format, load checkpoint with safetensors instead of PyTorch.",
			`)`
			`parser.add_argument(`
			`"--to_safetensors",`
			`action="store_true",`
			`help="Whether to store pipeline in safetensors format or not.",`
			`)`
Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`parser.add_argument("--dump_path", default=None, type=str, required=True, help="Path to the output model.")`
Device to use (e.g. cpu, cuda:0, cuda:1, etc.) (#1844) * Device to use (e.g. cpu, cuda:0, cuda:1, etc.) * "cuda" if torch.cuda.is_available() else "cpu" 2022-12-27 06:42:56 -07:00			`parser.add_argument("--device", type=str, help="Device to use (e.g. cpu, cuda:0, cuda:1, etc.)")`
unCLIP variant (#2297) * pipeline_variant * Add docs for when clip_stats_path is specified * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip_img2img.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip_img2img.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * prepare_latents # Copied from re: @patrickvonplaten * NoiseAugmentor->ImageNormalizer * stable_unclip_prior default to None re: @patrickvonplaten * prepare_prior_extra_step_kwargs * prior denoising scale model input * {DDIM,DDPM}Scheduler -> KarrasDiffusionSchedulers re: @patrickvonplaten * docs * Update docs/source/en/api/pipelines/stable_unclip.mdx Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> --------- Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-14 12:28:57 -07:00			`parser.add_argument(`
			`"--stable_unclip",`
			`type=str,`
			`default=None,`
			`required=False,`
			`help="Set if this is a stable unCLIP model. One of 'txt2img' or 'img2img'.",`
			`)`
			`parser.add_argument(`
			`"--stable_unclip_prior",`
			`type=str,`
			`default=None,`
			`required=False,`
			help="Set if this is a stable unCLIP txt2img model. Selects which prior to use. If `--stable_unclip` is set to `txt2img`, the karlo prior (https://huggingface.co/kakaobrain/karlo-v1-alpha/tree/main/prior) is selected by default.",
			`)`
			`parser.add_argument(`
			`"--clip_stats_path",`
			`type=str,`
			help="Path to the clip stats file. Only required if the stable unclip model's config specifies `model.params.noise_aug_config.params.clip_stats_path`.",
			`required=False,`
			`)`
Stable diffusion text2img conversion script. (#154) * begin text2img conversion script * add fn to convert config * create config if not provided * update imports and use UNet2DConditionModel * fix imports, layer names * fix unet coversion * add function to convert VAE * fix vae conversion * update main * create text model * update config creating logic for unet * fix config creation * update script to create and save pipeline * remove unused imports * fix checkpoint loading * better name * save progress * finish * up * up Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2022-09-15 16:07:32 -06:00			`args = parser.parse_args()`

Module-ise "original stable diffusion to diffusers" conversion script (#2019) * convert __main__ to a function call and call it * add missing type hint * make style check pass * move loading to src/diffusers Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-01-20 09:30:44 -07:00			`pipe = load_pipeline_from_original_stable_diffusion_ckpt(`
			`checkpoint_path=args.checkpoint_path,`
			`original_config_file=args.original_config_file,`
			`image_size=args.image_size,`
			`prediction_type=args.prediction_type,`
			`model_type=args.pipeline_type,`
			`extract_ema=args.extract_ema,`
			`scheduler_type=args.scheduler_type,`
			`num_in_channels=args.num_in_channels,`
			`upcast_attention=args.upcast_attention,`
			`from_safetensors=args.from_safetensors,`
misc fixes (#2282) Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-08 10:02:42 -07:00			`device=args.device,`
unCLIP variant (#2297) * pipeline_variant * Add docs for when clip_stats_path is specified * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip_img2img.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/diffusers/pipelines/stable_diffusion/pipeline_stable_unclip_img2img.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * prepare_latents # Copied from re: @patrickvonplaten * NoiseAugmentor->ImageNormalizer * stable_unclip_prior default to None re: @patrickvonplaten * prepare_prior_extra_step_kwargs * prior denoising scale model input * {DDIM,DDPM}Scheduler -> KarrasDiffusionSchedulers re: @patrickvonplaten * docs * Update docs/source/en/api/pipelines/stable_unclip.mdx Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> --------- Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-02-14 12:28:57 -07:00			`stable_unclip=args.stable_unclip,`
			`stable_unclip_prior=args.stable_unclip_prior,`
			`clip_stats_path=args.clip_stats_path,`
Module-ise "original stable diffusion to diffusers" conversion script (#2019) * convert __main__ to a function call and call it * add missing type hint * make style check pass * move loading to src/diffusers Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> 2023-01-20 09:30:44 -07:00			`)`
			`pipe.save_pretrained(args.dump_path, safe_serialization=args.to_safetensors)`