HNHacker News
TopNewBestAskShowJobs

standfest

188 karma · joined October 14, 2022

submissionscomments
standfest··on Fuck it, make it anyway
This attitude is pure punk! I couldn't agree more. The reason is statistics: even if LLMs can do a lot, they still rather interpolate than extrapolate. They do not excel, they average out. They deliver what is anticipated. So if you love a problem, if you have an intrinsic drive to solve something (like many of us do) - chances are high that you outperform what an LLM could achieve. I also believe as humans we shouldn't have to replicate redundant solutions but rather focus on deep dives, excel, push boundaries. And this is simply done by doing so for the sake of doing. This is art. So thank you for this statement, it made my day!
standfest··on ChatGPT Now Enforces Data Access Controls
update: if I log in via incognito mode, the toggle is set to "improve the model for everyone"=Off. So I guess somebody deployed a bug in the frontend.
standfest··on ChatGPT Now Enforces Data Access Controls
Got the same thing, though I sent a 'do not train on my data's request for my pro account. But the toggle in the settings panel cannot be deactivated. Is this a bug or a serious policy change?
standfest··on Show HN: Pipeline and datasets for data-centric AI on real-world floor plans
Agree that “data realism” is the quiet differentiator in mature visual generation domains.

Floor plans / technical drawings feel a lot less mature though — we don’t really have generators that are “good” in the sense that they preserve the constraints that matter (scale, closure, topology, entrances, unit stats, cross-floor consistency, etc.). A lot of outputs can look plausible but fall apart the moment you treat them as geometry for downstream tasks.

That’s why I’ve been pushing the idea that simplistic generators are kind of doomed without a context graph (spatial topology + semantics + building/unit/site constraints, ideally with environmental context). Otherwise you’re generating pretty pictures, not usable plans.

Also: I’m a bit surprised how few researchers have used these datasets for basic EDA. Even before training anything, there’s a ton of value in just mapping distributions, correlations, biases, and failure modes. Feels like we’re skipping the “understand the data” step far too often.

standfest··on Show HN: Pipeline and datasets for data-centric AI on real-world floor plans
Totally agree that for floor plans the bottleneck is usually label/geometry quality, not model architecture. We looked at CV early on, but real plan archives are a pretty adversarial input: ~100-year-old drawings mixed with modern exports, lots of drafting styles/implicit ontologies, low-res scans + distortion, and sometimes multiple conflicting “truths” for the same plan (revisions, partial updates, different sources). Even with decent models, you still pay heavily in expert cleanup.

So we optimized against the real baseline: manual CAD-style annotation. The “data-centric” work for us was making manual annotation cheap and auditable: limited ontology, a web editor that enforces structure (scale normalization, closed rooms, openings must attach to walls, etc.), plus hard QA gates against external numeric truth (client index / measured areas, room counts). Typical QA tolerance is ~3%; in Swiss Dwellings we report median area deviation <1.2% with a hard max of 5%. Once we could hit those bounds at <~1/10th the prevailing manual cost, CV stopped being a clear value add for this stage.

On ambiguity (doors vs windows, stairs vs ramps): we try not to “guess” — we push it into constraints + consistency checks (attachment to walls, adjacency, unit connectivity, cross-floor consistency) and flag conflicts for review. On generalization: I don’t think this is zero-shot across styles; the goal is bounded adaptation (stable primitives + QA gates, small mapping/rules layer changes). Trade-off is less expressiveness, but for geometry-sensitive downstream tasks small errors compound fast.

standfest··on Tim Cook and Sundar Pichai are cowards
I still hope for EU regulations at least
standfest··on Strategic Wealth Accumulation Under Transformative AI Expectations
i am currently working on a paper in this field, focusing on the capitalisation of expertise (analogue to marx) in the dynamics of cultural industry (adorno, horkheimer). it integrates the theories of piketty and luhmann. it is rather theoretical, with a focus on the european theories (instead of adorno you could theoretically also reference chomsky). is this something you would be interested in? i can share the link of course
standfest··on I wag, therefore I am: the philosophy of dogs
As a humorous reference I recommend to watch this rather old tv commercial: https://youtu.be/ExnWRuniooY?feature=shared
standfest··on Ask HN: How to change jobs with almost no interviewing experience?
Here are my 2¢: Leverage platforms that offer mock technical interviews (e.g., Pramp, Interviewing.io, probably there are others too). This approach lets you simulate the interview experience in a risk-free environment, getting you accustomed to the format and the pressure. It’s crucial to receive feedback, and these platforms pair you with industry professionals who can provide just that. This method is effective because it targets your interview skills directly, allows for rapid iteration based on feedback, and builds your confidence in a more controlled setting than actual interviews.

Aside from just technical skills, these mock interviews can help you articulate your thought process clearly, which is often as important as the solutions themselves in ML roles. Remember, it’s not just about getting the right answer, but also showing how you approach problems.

A side note based on a pattern I've observed so far: candidates who practice like this tend to perform better not just in technical assessments but also in explaining their past projects and teamwork experiences, which are equally critical parts of the interview process.

Hope this helps. Dive in, get that feedback, and refine your approach. Good luck!

standfest··on Elon Musk sues Sam Altman, Greg Brockman, and OpenAI [pdf]
i think this is the logical next step of a feud which only recently re-gained momentum two weeks ago https://www.forbes.com/sites/roberthart/2024/02/16/musk-reig...
standfest··on Elon Musk sues Sam Altman, Greg Brockman, and OpenAI [pdf]
i think this is the logical next step of a feud which only recently re-gained momentum two weeks ago https://www.forbes.com/sites/roberthart/2024/02/16/musk-reig...
standfest··on Ask HN: What GEN AI tools you built
atm mostly RAG setups for corporations, and a little bit in the compliance domain. not super exciting. last year at iccv paris, we did something correlated to HouseGAN and presented a new dataset to train such setups: https://data.4tu.nl/datasets/e1d89cb5-6872-48fc-be63-aadd687... and https://zenodo.org/records/7788422
standfest··on The impact of founder personalities on startup success
The study models the difference between the character traits of founders and employees. I wonder if any VC is already using similar methods.
standfest··on NetworkX – Network Analysis in Python
I usually write small functions for postprocessing the proposed layouts from the default algorithms. So far this was always more than sufficient.
standfest··on Gilles Deleuze – What is Philosophy? [audio]
I would like to support this observation. Obviously, they generated some language of their own, with lots of implications and references. Once you start reading it, it becomes a rabbit hole (back in the days my entry point was Adorno and Horkheimer). Words might be familiar, but their meaning is different. There was a reason to study, and the trend to render science accessible to laypeople was not yet born. Maybe even not wanted, to quote the ideas behind the concept of cultural industrial complex.
standfest··on Gilles Deleuze – What is Philosophy? [audio]
I am not a big fan of the Sokal Hoax and the follow ups. They made cheap money out of obvious misinterpretations, and did much more harm than anything else. While Chomsky, maybe from a US perspective, found some value, I am more with Derrida who analysed it as what it was: sad. I recommend his perspective https://philpapers.org/rec/DERPM
standfest··on Cisco ACI Preferred group, a pinch of inter-VRF leaking and L3Out
i like how you narrowed this defect down from a mysterious restriction in a white paper to a single line in the zoning table.
standfest··on [dead]
Hi! Thanks for sharing. I am the OP, we wanted to post this tomorrow. But I think that's now not needed anymore. If someone has questions, i try to answer here too.
standfest··on First trained open source floor plan recognition network
A trained network to annotate raster images of floor plans was just published by Archilyse under AGPL. You can try it under the link above, additional information is shared in the discord channel.

Previously, Archilyse has published the largest curated dataset of dwellings https://doi.org/10.5281/zenodo.7070951 and its full codebase https://github.com/Archilyse for data labelling, quality assurance, georeferencing , simulation, and benchmarking.

Archilyse is one of the world's leading R&D teams regarding the document class floor plan and cooperates with several famous universities.

standfest··on Largest open dataset of apartment models ever got published
matthias here. damn ninjas are cutting onions again. thanks a lot for your kind words, it was an incredible team effort over the last years to create this. we simply hope that the community is going crazy with the data. and there is one more, even bigger thing, we will announce soon. so if this data amazed you, buckle up!
standfest··on Largest open dataset of apartment models ever got published
Archilyse provides a SaaS tool to convert floor plans (raster images) into 3d models of buildings (IFC or GeoJSON) with embedded contextual information. the 3d model accuracy gets verified via additional data sources (governmental building hull data, client database entries) and all buildings are geolocated. subsequently, in a 25cm grid multiple simulations (3d view shed analysis, daylight, traffic noise, centrality, discrete metrics) are computed and aggregated into feature vectors. these are used for training AVMs (reducing prediction errors in half), architectural analysis (judging architecture competitions), construction cost estimation, life cycle cost estimation, energy load peak prediction, ... different research groups use this dataset to derive design patterns and to come up with augmented ai workflows for architectural design (like copilot) or as benchmarks for novel fitness/cost functions in MCO strategies.