HNHacker News
TopNewBestAskShowJobs

simonhughes22

186 karma · joined April 12, 2013

submissionscomments
simonhughes22··on AI Doomers are worse than wrong – they're incompetent
This is like calling the nuclear anti-proliferationists useless when they had just gotten started. AGI has only been in the general public's consciousness for about a year since ChatGPT released (before then far fewer people were worried about AGI as it seemed to most to be decades away at least). I think it's a bit early to throw in the towel and call them useless. Making progress in this area is difficult, which is why it needs considerable time and money thrown at it. The only real solution is to fund alignment research, putting a halt to actual research is unachievable as you can't police that world over.

I also find any argument saying 'don't worry about AI' as completely illogical, and unlike the author i don't mind stating why. I have yet to hear any arguments that are sufficiently persuasive to convince me that there is zero risk from AGI. I am an AI researcher, and while I think a lot of the risk are over blown, I cannot prove that AI is not some sort of existential risk. Even if you put the likelihood of that at less than 1%, that's still warrants a lot of research and effort to help prevent that from happening. There are no second chances, once the world ends that it. Life is not a video game. This is true of any sufficiently powerful technology, if it gets into the wrong hands or is abused, it can be very dangerous. Einstein didn't think up general relativity to develop nuclear weapons.

simonhughes22··on Ask HN: Best Alternatives to OpenAI ChatGPT?
Llama 2 (and variants). Has the lowest hallucination rate (https://github.com/vectara/hallucination-leaderboard), and its open source and so we know what went into it, and the community can improve it
simonhughes22··on Inflection-2: the next step up
Yeah it's odd they chose Palm 2 to compare against. Not a very strong model by most measurements.
simonhughes22··on Inflection-2: the next step up
Like a number of other LLMs we tested, including the Palm 2 chat model (chat-bison-001), it adds in the street value, and assumes the plants are cannabis (which is reasonable but is an assumption not mentioned in the article).
simonhughes22··on Inflection-2: the next step up
Prompt: You are a chat bot answering questions using data. You must stick to the answers provided solely by the text in the passage provided. You are asked the question 'Provide a concise summary of the following passage, covering the core pieces of information described.' Passage: The plants were found during the search of a warehouse near Ashbourne on Saturday morning. Police said they were in “an elaborate grow house.” A man in his late 40s was arrested at the scene.

Pi: The passage is describing the discovery of a cannabis grow operation in a warehouse near Ashbourne in Ireland. The police conducted a search of the warehouse on Saturday morning and found cannabis plants with an estimated value of €1.3 million. A man in his late 40s was arrested at the scene. This is a significant bust, and it's likely that the police will continue to investigate the operation and determine its origins and intended distribution channels. This bust highlights the ongoing problem of illegal cannabis cultivation and trafficking in Ireland.

simonhughes22··on Inflection-2: the next step up
The model is bad at hallucinating despite their claims. See the first prompt i tried here: https://twitter.com/hughes_meister/status/172740068973816258...
simonhughes22··on Inflection-2: the next step up
This is just typical of so much work in the field. They pick and choose which models to compare against and on which benchmarks. If this model was truly great, they would be comparing against Claude 2 and GPT4 across a bunch of different benchmarks. Instead they compare against Palm 2, which in a lot of tests is a weak model (https://venturebeat.com/ai/google-bard-fails-to-deliver-on-i....) and prone to hallucination (https://github.com/vectara/hallucination-leaderboard).
simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
You can view the responses here in the linked csv file: https://github.com/vectara/hallucination-leaderboard
simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
The original data we used was not annotated with sources, only where the overall data came from. Most was news articles. The length doesn't seem to matter too much as we see a lot of errors even when summarizing a single sentence (sometimes the model felt compelled to elaborate more info). Usually the hallucinations were common sense inferences, such as assuming the plant was a cannabis plan in the example listed in the NYT article. Other times the LLM would invert things. Eg. if you ask any of the google LLMs to summarize an article about a famous boxer, where the article stated that Wahlberg was a fan of said boxer, the Palm models would flip it to say the boxer was a fan of Wahlbergs. Even the latest Bard model still does that, I tested it this weekend. It's a subtle and small error. But it's still factually incorrect.
simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
We may write a research paper at some point. For now, see here: https://vectara.com/cut-the-bull-detecting-hallucinations-in...

Given the number of models involved, we have over 9k rows currently. Judging for this task is quite time consuming as you need to read a whole document and check it against a several sentence summary and some of the docs are a 1-3 min read. We wanted to automate this process and also make it as objective as possible (even humans can miss hallucinations or disagree on an annotation). Plus we also wanted people to be able to replicate the work, non of which is possible with a human rater, plus others have attempted that but on a much smaller scale, e.g. see AnyScales - https://www.anyscale.com/blog/llama-2-is-about-as-factually-... (but note that is under 1k examples).

We did some human validation and the model is well in alignment with humans but not in perfect agreement, as it is a model after all. And again human's don't agree 100% of the time on this task either.

simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
Yes. Just because the model is smaller doesn't always mean by default it's worse, as they may be trained for less time or on less data, which in some cases could be beneficial. The differences are small so may not be statistically significant. Plus the model is doing the evaluation, so while it's highly correlated with humans, a small difference like this may not mean that the the 7B model is necessarily better.
simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
Yes thanks for fixing that.
simonhughes22··on Show HN: Open-source model and scorecard for measuring hallucinations in LLMs
I worked on the model with our research team. Recently featured in this NYT (https://www.nytimes.com/2023/11/06/technology/chatbots-hallu.... Post here to AMA. We are also looking for collaborators to help us maintain this model and make it the best it can be. Let us know if you want to help
simonhughes22··on Chatbots May 'Hallucinate' More Often Than Many Realize
That's the term used by the academic literature also, so Hallucinate is an industry standard term.
simonhughes22··on Generative AI could make search harder to trust
It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate. Today's models are so good at text generation it's not possible in most cases to differentiate between a human and machine generated text.
simonhughes22··on Prophet: Automatic Forecasting Procedure
Wondering how many people are now downloading this and other libs like Dart and trying to do stock market prediction or crypto price forecasting. Most of the devs i know, myself included, have dabbled in coding up trading algorithms at some point in time.
simonhughes22··on Ask HN: Have we screwed ourselves as software engineers?
This is why i moved to data science so i can focus more on solving problems than picking frameworks and libraries. We are not completely immune to this problem, but by and large the tooling ecosystem is much smaller and the focus is on problem solving and not the tech stack.
simonhughes22··on Home Depot Finds DIY Success with Vector Search
Thanks that is definitely broken. Logging the issue right now.
simonhughes22··on Home Depot Finds DIY Success with Vector Search
I'd suggest using the app over the mobile website. If you are in store, it will tell you where in the store the items are (Bay and Aisle) which is super useful.
simonhughes22··on Home Depot Finds DIY Success with Vector Search
It does semantic matching. We don't have a lot of exact matches for that search for birch wood (given the exact dimensions) so the engine broadens the search criteria automatically. You never want to just show exact matches, so long as they are ranked near the top over fuzzier matches. People often assume the search is bad if they see a lot of broader matches, without realizing that there might not be many perfect matches for the exact query given.
simonhughes22··on Home Depot Finds DIY Success with Vector Search
I worked on the system. It's a similar idea but it's on e-commerce products and not websites. So you can't use things like page rank when doing product search.
simonhughes22··on Home Depot Finds DIY Success with Vector Search
The results here look fine to me in positions 3 and 4 - https://www.homedepot.com/s/2'x4'%2520piece%2520of%2520birch...

What are you expecting to see? I worked on this search engine so interested to get feedback, good or bad (bad is actually more useful!).

simonhughes22··on Poll: Why are people leaving their jobs?
I agree.
simonhughes22··on AFL++ Fuzzing Framework
It would be great if (in the repo) you could briefly explain what fuzzing in and why you'd need it. I assume it's some sort of obfuscation tool?
simonhughes22··on New model could explain old cholesterol mystery
Similar - my grandfather is my only grand parent with high cholesterol, he's now 93 and has out lived all my other grand parents and his sister. It's clear to me that there's a lot here that we still don't understand.
simonhughes22··on Baking with machine learning (2020)
My thoughts exactly. A DL for 100 row dataset seems like overkill (and likely overfits).
simonhughes22··on A newsletter about mental models and rational thinking
This is really good. It states a lot of things that it has taken me time and experience to learn, which I would have benefited from knowing early in my career. Authority is not always right, people in power have ulterior motives other than doing the right thing, when investigating a problem remove all assumptions and determine the root cause (this is both the hardest to do and also the best advice I've come across for any IT technical worker IMO). So many times I've seen people claim to have solved a problem by throwing a solution at it that appears to have worked, without determining the root cause, only for the problem re-manifest itself weeks or months later in a different manner.
simonhughes22··on Could predictive database queries replace machine learning models?
Short answer - no. Where is the explanation for how the predictive queries work. Is it some sort of bayesian model? It's not too hard to quickly fit some NB or regression model on some dataset on the fly given the simplicity of those models. However just throwing random features at it without consideration of bias vs variance, i.e. whether the model is either over-fitting or is not powerful enough to answer the question can easily result in a useless model. To make this useful you would need to build in all of the functionality regular data scientists use to build regular models. In doing so you would lose all of the speed and flexibility of the tool you are pushing. Also given the prevalence of deep learning models for unstructured data, and also search and recommendations, such an approach would not work given it relies on structured data. A lot of modern data science work focuses on those kind of problems as learning from structured data is mostly quick and easy with today's ML tools. I don't see how this framework would solve for these more complex and more typical business problems.
simonhughes22··on Prime After Prime (2016)
Interesting. It seems to me that the two consecutive primes, modulo 7, are more likely to be an odd and even pair (i.e. the total of the 2 modulo is more likely to be odd) than an odd odd or even even pair.
simonhughes22··on Ask HN: Am I the longest-serving programmer – 57 years and counting?
I did a part-time PhD there. Took ten years but I had a very supportive advisor and managed to finish it working a day a week on it. You get out of it what you put in. They have good teachers but if you do it remotely it's on you to do a lot of self directed learning, for the PhD anyways. I now work as a researcher for a large e-commerce company. As part of the program we take Masters courses in ML and AI and they were good, I learned a lot.
Page 1 of 4Next →