4,020 karma · joined September 30, 2014
Seems like there's a more concrete user guide here: https://docs.pola.rs/user-guide/gpu-support/#whats-supported...
It seems like they've done a great job of supporting most ops on the GPU already. The main one that is a really bummer is the categorical datatypes. Those are so common in ML workflows and using the native categorical datatype saves a lot of boiler plate as well as prevents you from messing something up during training!
I would would consider using a larger model for demonstrating inference performance as I have 7B models deployed on CPU at work, but GPU is still important training BERT size models.
Also, for my ML workloads the most common bottleneck is GPU VRAM <-> RAM copies. Doesn't this dramatically increase latency? Or is it more like it increases latency on first data transfer, but as long as you dump everything into VRAM all at once at the beginning you're fine? I'd expect this wouldn't play super well with stuff like PyTorch data loaders, but would be curious to hear how you've faired when testing.
This line in particular puts a bad test in my mouth, because my $2k LG G2 OLED has the worst support for casting I've ever experienced. In fact the software in general is so bad I was excited to pre-order the new Google TV Streamer this morning, so I don't have to deal with it again.
Really feels like if you need accelerated compute GCP is the better option these days. At least there you can rewrite in Jax if it comes down to it and opt for TPU's.
Also, the pricing is pretty shady. It seems like your GitHub PAT gives you free limited access, but if you ever want to move to a paid model, you have to transition to Azure. The shady part is that it's pretty hidden. You have to go to a specific model in the Marketplace, (for example: https://github.com/marketplace/models/azureml-mistral/Mistra...), then go to the bottom of the "Chapters" to the "Going beyond rate limits" section. There it just directs you straight to the Azure portal.
Edit: Nvm, I kept reading! Thanks for the interesting post!
In a major reversal, Google is ending a plan to eliminate cookies in its Chrome browser after four years of efforts, delays and disagreements with the advertising industry.
The decision to keep the pervasive tracking technology known as "cookies" in Chrome comes after a series of setbacks, as both digital-advertising companies and regulators objected to the plan and to Google's proposed replacement technologies.
Chrome users can already choose to block cookies in the browser's settings. Now, instead of eliminating them, Google will present users with a prompt to decide whether to turn cookies on or off, said the U.K. privacy regulator, which has been overseeing Google's plan to block cookies.
"We recognize this transition requires significant work by many participants and will have an impact on publishers, advertisers, and everyone involved in online advertising," Anthony Chavez, vice president of Google's Privacy Sandbox, the company's initiative to replace cookies, wrote in a blog post Monday. "In light of this, we are proposing an updated approach that elevates user choice...We're discussing this new path with regulators, and will engage with the industry as we roll this out."
Google first announced the plan to kill cookies in 2020, saying it would do so within two years to help protect users' privacy when surfing the linternet. Advertisers objected, saying that Google's plan to replace cookies would force them to shift spending to the search giant's digital-ad products.
In 2021, U.K. regulators opened an investigation into whether the plan would hurt competition in digital advertising. Google pledged to collaborate with the regulator and committed to give the agency at least 60 days notice before removing cookies to review any plan, and potentially impose changes to it.
As that investigation dragged on, Google's schedule to kill cookies by 2022 slipped.
In April, The Wall Street Journal reported that the British government's Information Commissioner's Office would issue a report criticizing Google's proposed replacement technologies as deeply flawed. A few days later Google said it would delay cookies' demise beyond last announced target date of the end of this year.
But I was always looking for other options for rebuilding the service within those constraints and found Quickwit when it was under active development. I really admire their work ethic and their engineering. Beautifully simple software that tends to Just Work™. It's also one of the first projects that made me really understand people's appreciation for Rust as well outside of just loving Cargo.
This actually looks really cool!
What I'm really excited about is a tool like Elicit using the new Google Gemini 1.5 Pro/Ultra models with the 2 Million token window sizes, filtering down the papers using traditional search and high quality meta-data, then critically, prompt/activation caching to make the tool economically viable.
Maybe it won't work better, but I'm willing to bet it'll find those really specific ideas/needles in the haystack a lot more often than vanilla RAG will.
My org would never allow that as we're in a highly regulated and security conscious space.
Totally agree about the BQ costs. The free tier is great and I think pretty generous, but if you're not very careful with enforcing table creation only with partitioning and clustering as much as possible, and don't enforce some training for devs on how to deal with columnar DB's if they're not familiar, the bills can get pretty crazy quickly.
https://cloud.google.com/bigquery/docs/geospatial-data#parti...
My org has a very large AWS spend and we got to have a chat with some of their SWE's that work on the geospatial processing features for Redshift and Athena. We described what we needed and they said our only option was to aggregate the data first or drop the offending rows. Obviously we're not interested in compromising our work just to use a specific tool, so we opted for better tools.
The crux of the issue was that the large problem column was the geometry itself. Specifically, MultiPolygon. You need to use the geometry datatype for this[1]. However, our MultiPolygon column was 10's to 100's of MB's. Well outside the max size for the Super datatype from what I can tell as it looks like that's 16 MB.
[1]: https://docs.aws.amazon.com/redshift/latest/dg/GeometryType-...