https://www.bloomberg.com/news/features/2022-06-10/how-citie...
5,971 karma · joined February 23, 2007
https://www.bloomberg.com/news/features/2022-06-10/how-citie...
We’re social creatures, chatbots already act as friends and advisors for many people.
Seems like a pretty good vector for a social attack.
Nothing wrong with setting out to build a Google instead of a Basecamp. They’re different kinds of company.
If anything it’s easy to underestimate the risk of building a low-risk business. They’re all hard.
Kudos to Konfig!
From https://arxiv.org/html/2409.11340v1
> Unlike popular diffusion models, OmniGen features a very concise structure, comprising only two main components: a VAE and a transformer model, without any additional encoders.
> OmniGen supports arbitrarily interleaved text and image inputs as conditions to guide image generation, rather than text-only or image-only conditions.
> Additionally, we incorporate several classic computer vision tasks such as human pose estimation, edge detection, and image deblurring, thereby extending the model’s capability boundaries and enhancing its proficiency in complex image generation tasks.
This enables prompts for edits like: "|image_1| Put a smile face on the note." or "The canny edge of the generated picture should look like: |image_1|"
> To train a robust unified model, we construct the first large-scale unified image generation dataset X2I, which unifies various tasks into one format.
Take for example: "A dog says \"Woof!\""
With a grammar, you’ll end up with "A dog says " when the model forgets to escape.
Which is valid JSON, but not what the model intended.
So it’s usually better to catch the exception and ask the model to try again.
Unless you’ve come across a sampler with backtracking? That would be cool
It’s pretty adept at most natural language tasks (“summarize this”) and performance on iPhone is usable. It’s even decent at tool once you get the chat template right.
But it struggles with json and html syntax (correctly escaping characters), and isn’t great at planning, which makes it a bad fit for most agenetic uses.
My plan was to let llama communicate with more advanced AI’s, using natural language to offload tool use to them, but very quickly llama goes rogue and starts doing things you didn’t ask it to, like trying to delete data.
Still - the progress Meta has made here is incredible and it seems we’ll have capable on-device agents in the next generation or two.
By way of example, 00-99 is 10^2 = 100
So, no, not the largest site on the web :)
The results are impressive - Llama 3 8b performs almost on par with GPT-4o across a wide range of tasks, not just logic and math.
Interestingly, the post-training process significantly improves model performance even without “thoughts” (the “direct baseline” case in the paper).
This isn’t a huge issue now, but with another 12 months of local AI progress it’ll leave a competitive opening for device-makers who ship more memory. (See Meta’s TPO paper for significant improvements to Llama, pending release.)
I skipped this iPhone cycle despite being on Apple’s iPhone upgrade program — a memory bump to the Pro line would’ve easily justified an upgrade.
I don’t think it’s fair to claim the weights are available if you need to hammer out a custom agreement with mistral’s sales team first.
If they had a self-serve process, or some sort of shink-wrapped deal up to say 500k users, that would be great. But bespoke contracts are rarely cheap or easy to get. This comes from my experience building a bunch of custom infra for Flux1-dev, only to find I wasn’t big enough for a custom agreement, because, duh, the service doesn’t exist yet. Mistral is not BFL, but sales teams don’t like speculating on usage numbers for a product that hasn’t been released yet. Which is a bummer considering most innovation happens at a small scale initially.
It's almost a pocketbook form-factor. I overlooked it initially because who wants a basic model? but the only thing I miss in practice is waterproofing. That, and the Oasis OEM cover which was unexpectedly nice, like a leather-bound pocketbook.
I’m not opposed to licensing but “email us for a license” is a bad sign for indie developers, in my experience.
8b weights are here https://huggingface.co/mistralai/Ministral-8B-Instruct-2410
Commercial entities aren’t permitted to use or distribute 8b weights - from the agreement (which states research purposes only):
"Research Purposes": means any use of a Mistral Model, Derivative, or Output that is solely for (a) personal, scientific or academic research, and (b) for non-profit and non-commercial purposes, and not directly or indirectly connected to any commercial activities or business operations. For illustration purposes, Research Purposes does not include (1) any usage of the Mistral Model, Derivative or Output by individuals or contractors employed in or engaged by companies in the context of (a) their daily tasks, or (b) any activity (including but not limited to any testing or proof-of-concept) that is intended to generate revenue, nor (2) any Distribution by a commercial entity of the Mistral Model, Derivative or Output whether in return for payment or free of charge, in any medium or form, including but not limited to through a hosted or managed service (e.g. SaaS, cloud instances, etc.), or behind a software layer.
The model is stock llama, fine tuned with a set of long documents to encourage longer outputs.
Most of the action seems to happen in an agent.
I don’t think that’s a dealbreaker - as long as the device can actually act as a standalone computer.
But it can’t, at least not for developers. It’s much closer to an iPad that only runs 10% of the apps you want. Maybe this is sufficient if your job only requires communication, and not, like, actual work.
So you end up with an expensive, socially awkward accessory to your MBP, which quickly gets left at home, because it doesn’t really do anything better than your existing devices.
(The one use case AVP handles well: watching a movie, in bed, alone. Which is kinda bleak.)
Image
- Terminus Research Group, from bghira of SimpleTuner/diffusers https://discord.gg/cSmvcU9Me9
- AI Toolkit https://github.com/ostris/ai-toolkit https://discord.gg/VXmU2f5WEU
- Stable Diffusion https://discord.gg/stablediffusion
LLM
- LLamaIndex https://www.llamaindex.ai https://discord.com/invite/eN6D2HQ4aX
- Nous research https://discord.gg/nousresearch
- LangChain https://discord.gg/hMrfPpUk
Platforms
- Replicate https://discord.gg/replicate
But if you’re in a few discords and a bunch of subreddits, you’re doing it right.
The most interesting stuff happens in GitHub PR’s, but you have to know where to look. Kohya’s misnamed SD3 branch has a ton of good flux hints, for example. It’s also where furkan gets pretty much all his content, before it gets paywalled.
Unfortunately, unless you participate full-time it’s hard to follow along. But if you really dig in and learn to modify your tooling (Comfy, kohya etc), you’ll start to come across some really impressive people who are all self-taught, and very accessible.
It’s totally possible to work your way up to the frontier with a few months of hacking. (And disposable income for GPU time.)
And the overlap between image AI’s and LLM’s is actually pretty great since they’re all transformers under the hood.
Civit, in my experience, is a good source for weights but most of the guides are written by people without much actual experience.
If you haven’t already, use tensorflow or wandb to get an intuitive understanding of your training parameters. It’s very easy to connect your tools to these services. This is by far the most helpful thing I’ve done, and something I really regret not doing sooner.
Heroku's price is a persistent annoyance for every startup that uses it.
Rebuilding Heroku's stack is an attractive problem (evidenced by the graveyard of Heroku clones on Github). There's a clear KPI ($), Salesforce's pricing feels wrong out of principle, and engineering is all about efficiency!
Unfortunately, it's also an iceberg problem. And while infrastructure is not "hard" in the comp-sci sense, custom infra always creates work when your time would be better spent elsewhere.
The universe is a big place; a decent writer can find a way to tell any story they want. (As demonstrated by the many IP-friendly reboots since 2004, when this was written.)
Peter Reinhardt from Segment has a must-watch talk for anyone interested in this topic https://youtu.be/_6pl5GG8RQ4?si=pogHC45L58U7K6mW
Sadly payment models are incompatible with how most people consume content – which is to read a small number of articles from a large number of sources.
I used AdBlock Plus for a while because it looked like the more popular ad-blocker.
I only uninstalled AdBlock Plus because it keeps displaying a "upgrade to premium" popup.
uBO has been such an improvement that now I worry I lost some geek gred for ever using ABP.