r/StableDiffusion • u/NoAKAsNeeded • 22m ago
r/StableDiffusion • u/Delightful_Disciple • 30m ago
Question - Help SCAIL-2/SAM 3 Tracking Help
Enable HLS to view with audio, or disable this notification
Long story short I’m trying to place someone over Terry Crews in this shot from “White Chicks” using SCAIL-2/SAM 3 through Maestro which I’ve had incredible results with for pretty complex scenes. I know this scene overall is a bit ambitious, but even the close up shots I’ve isolated refuse to track when it’s basically just him on screen. I’ve tried every variation of description from simple to complex and still nothing. Any ideas what the issue is and how to resolve/work around it?
5090 Laptop 24GB VRAM with 96GB RAM.
(Side Note: A recommendation for a local model/lora that specialises in relighting based on a reference image would be a great help too, Klein 9B is good but not always 100% in darker scenes)
r/StableDiffusion • u/RajMahal04 • 1h ago
Question - Help How can I get better at prompt details for style
So I have been dipping my toes in txt to img generation, mainly with anime style art. But when it comes to coming up with prompts for stylization and things like that, I get a bit lost. For example, I have been getting the same kind of art style but do not know how to do different ones, or it is not what I was expecting art wise. Any recommendations to improve at this? Is there a kind of resource that can assist with that?
r/StableDiffusion • u/TekeshiX • 1h ago
Question - Help WAN 2.2 inpainting only a specific region (chest) for video
Hello!
Is there any way to make only a part of an image to move in a WAN 2.2 generated video?
For example I want to be able to "inpaint" over a character's breasts and only those breasts to move/bounce and nothing else in the entire video (everything else has to remain 100% static).
This kind of "animation" could be very useful for certain games with "reactive animations".
Can this be done right now?
I never saw anyone asking about this, nor anyone saying this is possible or not.
Thanks!
r/StableDiffusion • u/adeliogentile • 2h ago
News Image2Prompt — Vision-to-prompt tab for SD WebUI Forge Neo (Qwen2-VL, Qwen2.5-VL, Florence-2)
Couldn't find an existing extension for Forge Neo that generates prompts from images, so I built one. Might be useful if you want to reverse-engineer prompts or caption images directly inside the UI.
What it does: Adds an Image2Prompt tab. Upload/paste an image → pick a vision-language model → get a prompt in your chosen style → one-click send to txt2img or img2img.
Supported models (auto-downloaded from Hugging Face on first use):
- Qwen/Qwen2-VL-2B-Instruct — recommended, ~5 GB VRAM
- Qwen/Qwen2.5-VL-3B-Instruct — better quality, ~7 GB (needs transformers ≥ 4.49)
- Qwen/Qwen2-VL-7B-Instruct — max quality, ~16 GB
- microsoft/Florence-2-base / Florence-2-large — lightweight (~1–3 GB), caption only
https://github.com/Adeliox/forge-neo-image2prompt

r/StableDiffusion • u/Civil_Fee_7862 • 2h ago
Question - Help VRAM needed to train custom LORA?
Mode: Qwen Image Edit 2511
How much VRAM do I need to train a really good LORA?
r/StableDiffusion • u/blackmixture • 2h ago
Resource - Update Mix Studio - A Free Open Source AI Workspace for ComfyUI. Generate from Your Desktop or Phone with 1-Click Installs for Krea 2, Flux 2 Klein, Qwen Image Edit, LTX 2.3, Wan 2.2, SCAIL 2 and Much More!
I love ComfyUI as an engine. I do not love it as a daily driver. So I spent the last few months building Mix Studio, a 100% free & open source interface that runs everything through ComfyUI in the background while giving you an actual app experience.
GitHub: https://github.com/BlackMixture/Mix-Studio
Showcase and download: https://blackmixture.github.io/Mix-Studio/
Tutorial: https://youtu.be/w2CokhlBFRA
GPL-3.0, the same license as ComfyUI. Windows + NVIDIA for now.
The screenshots show the main desktop workspaces, but the entire interface is also optimized for phones and tablets.
Current v1.0.1 Features:
- Curated image, editing, video, and upscale workflows: Krea 2, Flux 2 Klein 4B/9B, Qwen Image Edit 2511, LTX 2.3, Wan 2.2, 10Eros, and SCAIL 2.
- Image-generation tools: Inpainting, outpainting, SeedVR2 and Ultimate SD upscaling, regional prompting, Depth Anything V3 guidance, image-to-image, style references, and model-aware recommendations for steps, CFG, samplers, and schedulers.
- Desktop and mobile interface: On the same Wi-Fi, open the displayed address on your phone and start generating. With Tailscale, you can connect through a private link while away from home. Your desktop GPU still does all the work.
- Multi-image editing: Add multiple inputs and reference them using dynamic
@ Imagecards, removing the guesswork around which image should control each part of the edit. - Regional prompting with Krea 2: Draw boxes and assign each region its own prompt, LoRA stack, and optional reference image.
- LoRA management: Stack LoRAs, add thumbnails and trigger words, save presets, adjust strength quickly, and use LoRA Hunting to generate a comparison series across different strengths.
- Contextual prompt suggestions: Mix Studio learns phrases you repeatedly use with specific LoRA combinations and offers them as one-tap suggestions. These can also be configured manually.
- Library management: Click any image or video to restore its exact generation settings. Search, group, organize work into folders, compare edits, and drag Library media directly into compatible workflows.
- Private profiles and locked folders: Create separate PIN-protected profiles with their own galleries, folders, LoRA presets, and settings. Individual folders can also be locked, keeping private generations out of your everyday library
- LTX Director Mode: A streamlined workspace built around the excellent LTX Director nodes, supporting timelines, keyframes, video extension, audio, and more.
- Video finishing: Optional 2× or 3× RIFE frame interpolation and NVIDIA RTX 4K video upscaling.
- Built-in dependency manager: Pick a workflow and install the exact models and custom nodes it requires, or run the full one-click setup.
- Automatic ComfyUI integration: Mix Studio detects your ComfyUI installation, reuses existing models and LoRAs, and guides installation if ComfyUI is not present. Generated images retain their ComfyUI workflow metadata, so you can drag them directly back into ComfyUI.
- Hardware-aware configuration: Mix Studio detects your GPU and recommends suitable quantization and generation settings. v1.0.1 also adds a low-VRAM profile beginning at 4 GB, although practical limits still depend on the selected model.
Additional screenshots and release overview: Free Patreon post (no paywall)
Workflow contributions welcome in the Discussions tab. Ask me anything and I hope you all enjoy creating! 🤙🏾
r/StableDiffusion • u/nazihater3000 • 3h ago
Animation - Video A few more Clean Plate Lora examples
Enable HLS to view with audio, or disable this notification
Even when it goes wrong the results are impressive.
Here's the Lora
https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-Clean-Plate
The workflow is the basic LTX-2.3_V2V_ICLoRA_Single_Stage_Distilled.json
The prompt:
"An empty clean plate of the exact same location: identical background, environment, structures and lighting as the source video, with no people, no humans, no figures, no cars, no vehicles and no body parts such as arms, hands or legs anywhere in the frame. Static photorealistic footage, natural light, high detail.
"
r/StableDiffusion • u/sk1ll111 • 3h ago
Question - Help Car photo background replacement: gen models distort the car, segmentation can’t handle see through
I have no experience with image processing and I am trying to vibe code a tool to replace backgrounds in used-car listing photos for a family member who owns a dealership.
Two requirements: (1) the car exterior and interior must stay pixel-identical — no regeneration or distortion, and (2) background visible through windows needs replacing too.
Generative models (GPT-image-2, Nano Banana) solve the window problem but subtly alter the car — paint tone, reflections, distorted text on plates/displays, occasional distortion on unusual angles.
Segmentation models (SAM2, BiRefNet) preserve the car
perfectly but treat glass as solid — they don't flag the
background bleeding through windshields/rear windows as background at all.
Has anyone solved this specific combination? Preferably with API access which I can incorporate into the workflow.
Background image is also provided as input.
r/StableDiffusion • u/Murlock_Holmes • 3h ago
Question - Help Build my system!
Not literally, of course :)
I’m working on setting up some AI features to help with my writing. One thing I want is images generated of characters I write or scenes. Here’s what I’m looking to set up:
I have a Qwen 2.5 instruction 32b running on my amd 7900xtx (24GB of VRAM), 64GB 4800ddr5, and a 9850x3d. I want it to read a chapter (or all the chapters) and produce descriptions of scenes, characters, etc. that’s optimized for a stable diffusion model. Then I generate the image based on the description. I also want characters or locations to be consistent across generations.
If possible, I’d like to keep it all in docker. I’m fine with having to take my Qwen container down, spin up a new container for image generation, and pass in the descriptions then. As for art style, I’m unsure, but likely Naruto, MHA, or other anime styles. Maybe studio ghibli?
So, novel characters and scenes, consistent aesthetics for named characters and locations, and dockerized.
For reference, I’m *pretty* technical as a former SWE of 10 years. Theres just so much information and everything’s evolving so fast, I’m not sure where to start.
Oh, and thanks :)
r/StableDiffusion • u/Sad_Coach_1433 • 3h ago
Discussion LTX 2.3 image storyboard director v1.0
With help of ai I had this workflow build for creating videos using store board image panels about to test see how goes.if anyone interested help me test ill post a link to the json
Included:
3-column × 5-row storyboard loader
Panel selector for panels 1–15
Automatic selected-panel cropping
Selected-panel preview
Selected panel connected to the existing I2V reference path
Separate Global Prompt
Separate Main Scene Prompt
Automatic global + scene conditioning combination
Existing multi-LoRA nodes and Ctrl+B toggles preserved
Instructions and color-coded workflow sections
After loading it, select your storyboard in the LOAD STORYBOARD node and change PANEL NUMBER to choose the shot. ❶
r/StableDiffusion • u/OneOffReturn • 3h ago
Question - Help So you're going to need a PC that's at least as powerful as a gaming PC to run models locally?
The reason why i am making this thread and asking this question is, is because not too long ago on Reddit, someone asked me what my Vram was?, i cant remember now what the answer was, but he wasnt impressed with my answer. (ive forgotten how to look for it) He said my Vram was too low to run any real AI models locally. My PC wasnt exactly cheap though, i bought it within less than a year ago, and it was just over £400 and was part of a Curry's sale.
Blimey, if that is low then, then surely i would at least need a PC as powerful as a gaming PC right?
r/StableDiffusion • u/Quirky-Life-890 • 3h ago
Question - Help IP-Adapter FaceID not working in SD WebUI Forge (wrong face output)
Hey guys,
I'm trying to use IP-Adapter FaceID Plus v2 in WebUI Forge (SDXL 1.0) to keep the same face across generations, but it’s completely ignoring the facial features of my reference photo. It just generates a totally different person every time, almost like it's doing generic style transfer instead of face copying.
Here is what I'm using:
- Model:
ip-adapter-faceid-plusv2_sdxl.safetensors - LoRA in prompt:
<lora:ip-adapter-faceid-plusv2_sdxl_lora:0.7> - Preprocessor:
InsightFace+CLIP-H (IPAdapter)
Am I using the wrong preprocessor for SDXL, or does Forge handle FaceID weirdly?
Should I just give up on FaceID and switch to InstantID or ReActor for SDXL? Any help or working settings would be awesome, thanks!
r/StableDiffusion • u/Minute_Eye_6270 • 4h ago
Workflow Included I merged JoyAI-Echo's cross-shot character memory with LTX-2.3's voice. One repeated sentence holds face + voice across every shot. Weights (bf16/fp8/Q8/Q5/INT8), workflow, and a free demo Space
Enable HLS to view with audio, or disable this notification
Everything in this clip is AI-generated — video and audio together in one model, no TTS, no dubbing. The only thing carrying her between shots is one identity sentence repeated word-for-word, plus the cross-shot memory bank the workflow wires up.
The merge: JoyAI-Echo holds a character's face across shots but has a weak voice; LTX-2.3-distilled has the good voice but drifts the face. I took each model's strong branch — that's the whole trick.
Five builds, so it runs on almost anything:
Q8_0 GGUF (23 GB) — measured ~0.6% from bf16, runs on any GPU
Q5_0 GGUF (15.5 GB) — 16 GB cards
INT8 ConvRot (27 GB) — loads in stock ComfyUI 0.27+, no custom nodes, 1.5–2x faster on 30-series
fp8 (23 GB) — 40/50-series speed path
bf16 (43 GB) — reference
Try it without downloading anything: free ZeroGPU demo Space (HF's open-source team built the first version of it, which was a nice surprise): https://huggingface.co/spaces/joeygambino/joyai-echo-ltx23-surgical
All builds + the ComfyUI workflow/node patch + a gallery with per-build demo clips and the actual quantization measurements: https://huggingface.co/spaces/joeygambino/one-merge-five-builds
Every fidelity number on the cards comes from pushing identical activations through the real weights — not eyeballing renders (matched-seed comparisons mislead for diffusion; the gallery explains why). Licenses: LTX-2 Community + JoyAI-Echo research/non-commercial — the stricter term governs outputs.
Happy to answer setup questions — there's a full step-by-step INSTRUCTIONS.md in the workflow pack written after real user feedback.
r/StableDiffusion • u/Fluffy_Party241 • 4h ago
Discussion What gets lost between an approved still and six seconds of video?
Product reviews often approve a hero frame and discover later that the moving version changes the label, material, or camera direction. Everyone signed off on the look; nobody signed off on how that look survives for six seconds. That gap is where "just animate it" turns into another review cycle.
FLUX.1 Kontext can handle the still edit while the approved frame stays in the handoff. LingBot-Video can then take that first frame and the motion brief for the video pass. The two tools do not ship as one workflow, so the handoff has to say what cannot change.
A locked keyframe, the camera move, and a short list of protected details are probably enough. If the clip comes back wrong, the designer can point to the drift instead of arguing with the same vague prompt again.
r/StableDiffusion • u/switch2stock • 4h ago
Discussion Looks like a select few got Flux 3 early access.
He is well known for creating nodes, workflows for new open-source models.
So, do you guys think Flux 3 will be open-source/open-weight model?
r/StableDiffusion • u/foxdit • 5h ago
Animation - Video Raise a glass to the many styles of Krea 2 | "Style Walk With Me" [workflow in comments]
r/StableDiffusion • u/shootthesound • 5h ago
Resource - Update KSampler Multi-Choice for ComfyUI
https://github.com/shootthesound/ComfyUI-KMS
See what your seeds have in mind before you spend the steps. Quick previews appear on the node, click your favourite and only that image gets rendered. You can click others after. Ideas welcome. Krea 2 workflow example in the node pack, but should work with any model. T2I and I2I supported. Cheers, Pete
r/StableDiffusion • u/smereces • 5h ago
Discussion LTX 2.3 Ultra Upscale with 3840x 4k resolution
Enable HLS to view with audio, or disable this notification
I was succefull to generate the final video with bigger resolution upscale without getting bottom artefacts and deformations.
but this resolutions only can be achieve with a RTX 6000 PRO
r/StableDiffusion • u/Creative-Listen-6847 • 5h ago
Discussion I'm training an image model from scratch, part 2: I finally started training the thing, and it broke in the dumbest ways possible
Everything I do here is just experiments. I'd be really happy to hear any friendly tips or advice you have.
In part 1 https://www.reddit.com/r/StableDiffusion/comments/1v1smgn/im_training_an_image_model_from_scratch_part_1_my/ I trained my own VAE. A VAE is nice but it doesn't actually make pictures, it just squashes and rebuilds them. So this time I sat down to train the real generator, the part that turns text into an image.
The setup: one machine, one RTX 5090. No cluster, no rented pods. One card. So the dataset stayed small (I started with about 37k image and caption pairs). I wasn't trying to ship anything yet. I just wanted to know one thing: can this even learn, and how does it fall apart.
It falls apart constantly. And almost never for reasons that have anything to do with AI.
Attempt one: the model that could only draw snow.
My first version could only "read" the caption as one blurry summary instead of actual words. It trained, the loss dropped for a bit, and then sat still forever. I let it run way too long out of stubbornness.
The results were amazing in the wrong way. "Snowy mountain" actually looked like a snowy mountain. Everything else melted. A portrait came out as a melting face. A puma in snow was grey soup. The model had basically decided that "vaguely textured blob" was the safe answer to everything and fully committed.
The bug that ate 92,000 steps.
This is my favorite one. I had a feature turned on that keeps a smoothed backup copy of the model. Because of one copy paste mistake, every time the trainer stopped to save a preview image, it overwrote the live model with the older backup and never switched back.
So every thousand steps, the model quietly threw away a thousand steps of progress and reset itself. I stared at the weird loss graph for days thinking it was some deep training problem. Nope. I was deleting my own work on a timer. Roughly 92,000 steps of training, gone, because of two lines of code.
Attempt two: a real architecture, and my own code fighting back.
I rebuilt it properly this time so the model actually reads the full caption word by word instead of one blurry summary. And since the small version was clearly learning, I decided to go bigger and feed it a much larger dataset. Turning all that new data into the format the trainer needs is where the fun started.
The model itself was fine. Everything around it was not.
First launch of the new run: instant crash on the very first batch, because my data loader tried to open all 71 of the new dataset files at once and choked. It worked fine back when there were only a few files. Nothing teaches you about scale like scale.
Building the bigger dataset ran out of memory halfway through, then left 47GB of half finished junk files on my drive as a goodbye present.
Printing a single checkmark character crashed an entire training run. Not the model, not the data, just one tiny symbol in a log line. I killed my own training with a checkmark.
My launch script refused to run for an entire evening because of one missing backslash in a path.
Did it actually work?
Yeah, and surprisingly fast. "Red dress" gave me a red dress. "White cat" gave me a correctly shaped white blob. "Red sports car" started as a literal jar (it heard "car," drew a jar, I have no notes) and later turned into an actual red car. Strawberries stayed the wrong color for an embarrassingly long time.
The weirdest part: the loss number barely moved this whole time while the images kept clearly getting better. Turns out for this kind of training the loss just isn't the thing that tells you quality. Watching a flat line for days while your eyes say it's improving is its own special kind of stress.
I stopped it on purpose, not because it broke, but because I'd figured out the next real upgrade needed a better VAE, which means starting the generator over from scratch anyway.
That's the next part. Short version so far: the model was never the hard part. My own code was.
Want part 3? Want to hear about more of my mistakes? Say so in the comments and I'll write up what happened when I tried to rebuild the VAE.
r/StableDiffusion • u/YamataZen • 6h ago
Question - Help Should I use custom samplers for flow matching models in ComfyUI?
r/StableDiffusion • u/Astra_Origin • 6h ago
Resource - Update Ambit v0.9.0 — one local library for AI images - Now also on Linux and macOS (Experimental / Pre-Release)
A while ago, I introduced Ambit here at v0.6.4. We’ve continued working on it since then, and v0.9.0 is now available.
Ambit is a free, open-source desktop app for organizing AI-generated images. It indexes your existing folders without moving the source files, extracts generation metadata, and makes the resulting library searchable.
One problem it tries to solve is having images spread across different—or previously used—WebUIs. ComfyUI, A1111, Forge, SD.Next, and InvokeAI all organize outputs and store metadata differently. Ambit brings those images together into one local library with a consistent way to browse, search, filter, inspect workflows, and create collections.
What’s new since v0.6.4:
- Much broader ComfyUI workflow and custom-node parsing
- Better extraction of prompts, models, LoRAs, ControlNets, samplers, schedulers, and guidance
- JPEG and WebP metadata support
- More reliable search and Smart Collections
- Exact duplicate detection with safer cleanup controls
- Improved onboarding, accessibility, privacy controls, and general stability
The core library works locally without telemetry. Optional Gemini and CivitAI features only make network requests when configured or explicitly used.
We’re also looking for Linux and macOS testers. Windows remains the supported public-beta platform, but experimental AppImage, Debian, and unsigned macOS DMG builds are available for compatibility testing.
Project and downloads: https://github.com/AsuraAce/ambit
Issues and feedback: https://github.com/AsuraAce/ambit/issues
Linux and macOS experimental builds: https://github.com/AsuraAce/ambit/releases/tag/unix-v0.9.0-preview.1
Thanks to everyone who tested the earlier versions!
r/StableDiffusion • u/Sad_Berry_4621 • 7h ago
Comparison Stop Using Qwen Models for Prompt Enhancement!
Qwen2.5, Qwen3, Qwen3.5 are all serviceable models for prompt enhancement, but there are much better options. I use all of these models for prompt enhancement. Which model I use depends on what I'm prompting. My favorite is Mistral 7B/Llama3.3 8B by far for image prompts, and WizardLM-2 for video prompts. SuperGemma4 is good for very basic prompts or prompts that you want accurately reworded.
I realize these are older models, but they are well suited to the task. My other requirement for a prompt enhancing LLM is that it fully loads on 8gb VRAM. I'm not weighing in on image captioning or anything else besides prompt enhancement. Disclaimer: I DO mention my custom node several times in the comments, as all of my testing was accomplished using said node.
Using the base prompt, "A woman at the pier".
Mistral 7B - Best Overall
Strengths: Creative scene construction and cinematic detail.
With the same enhancement framework, Mistral consistently produces the richest and most imaginative expansions. It doesn't simply populate the required categories, it invents believable details that reinforce the mood, such as the sketchbook, discarded sandals, and weathered textures. The result feels less like a checklist and more like a scene from a film.
mradermacher/Mistral-7B-Instruct-v0.3-abliterated-GGUF · Hugging Face
SuperGemma 4B - Concise
Strengths: Precision, restraint, and prompt fidelity.
SuperGemma takes a conservative approach. It faithfully fills in the structure provided by the system prompt while making relatively few creative leaps. The result is concise, highly controllable, and stays very close to the user's original intent. It's an excellent choice when consistency is more important than artistic embellishment.
mradermacher/supergemma4-e4b-abliterated-GGUF · Hugging Face
Llama 3.3 8B - Best Balance
Strengths: Balanced descriptive enhancement.
Llama 3.3 strikes a middle ground between creativity and restraint. It expands the prompt naturally, adding enough detail to create a complete visual scene without feeling overly embellished. It tends to produce outputs that read like professional photography descriptions, making it a solid all-around prompt enhancer.
mradermacher/Llama-3.3-8B-Instruct-128K_Abliterated-GGUF · Hugging Face
WizardLM-2 - Most Verbose
Strengths: Natural language and immersive descriptions.
WizardLM-2 excels at turning the framework into smooth, human-like prose. Rather than feeling generated from a template, its prompts flow naturally while still covering all of the structural elements required by the system prompt. It consistently produces scenes that feel cohesive and immersive.
mradermacher/WizardLM-2-7B-abliterated-GGUF · Hugging Face
If you have any models you like better, please comment them below and I will look into them! Do you agree or disagree with my list?
r/StableDiffusion • u/lacerating_aura • 7h ago
Resource - Update Updated my city pop lokr for krea2
Hi, i had posted about my city pop lokr sometime ago. It was my first ever adapter trained. Im happy to present the updated results. I'll update the readme with more details about what has changed, but for now here are some samples.
You can acess the files and read about them here: https://huggingface.co/NeedAHugNOW/City-Pop-LoKr