HNHacker News
TopNewBestAskShowJobs

neilxm

141 karma · joined August 3, 2013

Founder at Placenote. We're building an A.I. assisted interior design tool
submissionscomments
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Yes! That's where our heads are at as well. The reality with a lot of multimodal / image proc style code is that it's never truly serverless - image manipulation in node.js is tragically bad so you always end up needing python endpoints to do it.

Re: version / client languages etc - right now we don't have block versioning but it's definitely going to be required. As of now the blocks are each their own endpoint, by design. We're thinking about allowing people to share their own blocks and perhaps even outsource compute to endpoint providers, while we focus on the orchstration laters.

Better observability and monitoring is definitely on the docket as well. Especially because some of these tasks take a really long time - some times even going past the expiry window of the REST api. We'll be switching over to queued jobs and webhooks

neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Thanks! We'll ponder over this one :)
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
YEP! If you've used blender you'll notice the parallels with shader nodes :)
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
It's a combination of things. The idea is that you can build workflows that chain functionality from ai models, as well as lower level image processing tasks. For lower level tasks we use the usual suspects - PIL, ImageMagik, OpenCV etc.
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
We're working on shareable graphs and premade recipes! I actually started sharing a few on our blog - here's an example: https://blog.mlblocks.com/p/auto-generate-banner-images-for-...

haha, the tilt animations are a by-product of my obsession with Trello. :)

neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Thats true. We started off with a base set of blocks but i think the real utility will come in the easy orchestration and api end point building. We're pushing in the direction of apis and shareable workflows so hopefully some of these comparisons get clarified soon
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
We started this to solve bulk processing issues we had when building a previous eCommerce tool so I 100% know what you mean. We're adding API support soon and we'll add some examples of how to connect this to Shopify or something like Airtable/ Strapi / Retool etc for workflow automations
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Thank you!!
neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Thanks you beat me to it :)

That being said, you're not wrong. It's definitely inspired by ComfyUI. But, with much simpler abstractions, much broader utility and extensions like building a user front end coming up shortly

neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
To add to pj's comment -

We are adding more blocks constantly. We're also considering allowing the community to push their own blocks using an open api schema.

neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Oh yea I know what you mean. There are several parallels here with shader nodes for sure. We've been thinking about a voyager/agent-style approach where an agent can start to learn "skills" where skills are individual blocks. Each skill represents a certain function applied to an image and based on a specific instruction set we should be able to craft a sequence of actions that will lead to that result.

One way to leverage that is building the graphs via a prompt, but another way might be to not think of the workflow as a pre-constructed graph at all. Rather perhaps we build dynamic graphs whenever you ask for a certain action - like a conversational image editing interface.

So you say something like make the woman's hair purple. We apply segmentation to the hair, and then add a puple color overlay exactly to that area.

neilxm··on Show HN: ML Blocks – Deploy multimodal AI workflows without code
Nice! Sharing workflows is coming up in approximately 2 sprints. We're working on 2 flavors of sharing. The first is sharing the workflow directly and letting someone copy it for the dev community. The more interesting option though is the second, where we'll let you build a read-only dashboard that will just show inputs and outputs. that should be useful when you share it with a marketing team that doesn't need to mess around with the graph but would use the workflow for things like repetitive image editing tasks.
neilxm··on Show HN: An AI agent that Auto-generates Illustrations for any Blog Post
Good ideas! Thanks! I'm definitely going to implement the style picking next
neilxm··on Show HN: Instant Banner Images for Any Blog Post Using GPT3.5 and SDXL
Currently there's no way to steer it, but I'm thinking about ways I can allow steering the image gen. Maybe just showing the prompt open ai spits out so you can edit it, but I wonder if there are better ways to allow control
neilxm··on Show HN: Instant Banner Images for Any Blog Post Using GPT3.5 and SDXL
BannerGPT is an AI agent that can auto-generate banner images for any blog post.

The premise is pretty simple. You paste your article into BannerGPT, it analyzes the text and generates a banner image that can complement your article.

Under the hood, the raw post is sent to GPT 3.5 with a prompt to describe a scene that illustrates the main ideas in the text. The description is then sent as a prompt to Stable Diffusion (hosted on Replicate). Finally the title is framed on the image and turned into a downloadable image.

neilxm··on Show HN: An AI agent that Auto-generates Illustrations for any Blog Post
BannerGPT is an AI agent that can auto-generate banner images for any blog post.

The premise is pretty simple. You paste your article into BannerGPT, it analyzes the text and generates a banner image that can complement your article.

Under the hood, the raw post is sent to GPT 3.5 with a prompt to describe a scene that illustrates the main ideas in the text. The description is then sent as a prompt to Stable Diffusion (hosted on Replicate). Finally the title is framed on the image with a nice gradient overlay and turned into a downloadable image.

If you have a personal or company blog, BannerGPT is a nifty little tool that can help add some color to your posts.

Here’s a hosted version to play with: https://bannergpt.dabble.so

I’m curious to see this tested in the wild, so if you share any generated banners in the comments or on twitter I can add them as examples on the site.

My twitter/X handle is @neilxm

neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
yea true. some trade offs. i'm going to look into consolidating everything into just the nextjs piece.
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
awesome ideas! the configurability and flexibility of 3D models is a huge advantage over a pure 2D approach these scenarios.
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
yea for sure. i guess i'd like this to be something usable on mobile as well. but directionally you're right. there's a lot to be gained from optimizing for local compute in many applications. something to look into.
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
Yea! Shadows on the item or shadows propogated back onto the scene
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
That’s the idea for next steps.

Basically if you generate a backdrop and then estimate light direction you can inverse render that onto the 3d model given all the depth information you get for free from the model

neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
So the api used for inpainting is very easily swappable to a local instance. But when you say consumer hardware, I think you might be overestimating how capable most people’s computing setups are.
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
Yes and no. So I did it because I honestly get annoyed with pixel operations in node js. The same thing in Python with the pillow library achieves the same things in quarter the lines of code as you would need in something like sharp in nodes

The other reason I used a Python backend was that I want to extend it to more involved image processing, like producing control net inputs or post processing the end result.

Does that make sense ? Open to better ideas though

neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
Hah. Yep!
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
bingo
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
Yea you don't absolutely need a 3D model. The benefit of 3D though, is you can define any camera angle and perspective of the main product. now of course you just take a picture of a product from any angle if it were in front of you, but that's not always feasible. the use case here is scalable photo generation for ecomm stores with thousands of inventory items.

Additionally, 3d means more than just camera angle control. you can define a scene in 3D and send it into control net to produce a very specific image

neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
So because this is a threejs canvas, we could build an interface to add lights anywhere, to adjust lighting on the model. this is an early version so it's not in there, but really easy to add in code. perhaps that will be the next feature.
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
thank you!
neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
It's not that different right now you're right about that. It chains that sequence of steps together and I'm not really sending any meta data about the objects pose to SD, I just leave it up to the model. But the next step here is leveraging more control-net approaches and thats when it starts to really take advantage of 3D more than just the 2D snapshot minus background.

For example, a 3D editor to allow for simple low poly style scene creation, that then serves as conditioning input to control net. For example staging a model of a chair in a sketchup model of a living room with super basic furniture elements in low poly 3D models. You pass that into SD and out comes a fully rendered image. At that point i think you could argue that stable diffusion could be used as a platform agnostic renderer like VRAY but for any 3D modelling tool.

This was my version 1 haha

neilxm··on Show HN: Generate Stable Diffusion scenes around 3D models
Reading your question again, I should clarify - it's not quite compositing a random background under the 2d image. it's using stable diffusion to "in-paint" a background that makes sense spatially, with correct shadows, lighting and perspective to match your prompt. That being said, there's still a lot to be done to deeply integrate the 3D and 2D spaces and that's just going to be part of the ongoing exploration in this project.
Page 1 of 2Next →