HNHacker News
TopNewBestAskShowJobs

simonmesmith

185 karma · joined August 13, 2022

submissionscomments
simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
I don’t know of any other models that support function calling like OpenAI’s, unfortunately.
simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
Great question, and something I’ve thought about.

The main issue right now is that I’m relying on the OpenAI API’s function-calling capabilities to enable the use of functions in workflows.

If we switch to other LLMs, we’ll need to create some additional wrapping around them to allow such function calling. (As far as I know. If any open source LLMs already have function calling built in, let me know.)

LangChain (which I’m not using due to personally finding it overcomplicated and too obfuscating) does have function calling, but it uses approaches like REACT (if I remember correctly) that aren’t as reliable as OpenAI’s approach.

simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
Thank you for saying so! It’s very gratifying that I’m not the only one who finds this useful :)
simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
I definitely plan on adding more functions, and hopefully having others do the same. And I think these should be versatile things like get_url(), which I added today, versus the very specialized plugins that seem to dominate something like ChatGPT’s plugin space. I think we want reliable, versatile building blocks as functions.
simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
Yeah, I saw the Magic Loop on HN and was like, “oh no, time to dump my project!”

But I think there are some key differences. I’m not sure if Magic Loop is open source, for example, so I don’t know how they’re currently building their workflows.

Automating API calls is neat, too, but I get a bit anxious with too much automation when you want to run a workflow with consistent results. My gut says you want to have the functions here pretty locked down and tested, rather than rely on automating API calls. But maybe I’ll be proven wrong.

simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
Perhaps we could give people the option of saving workflows in different places. I went with JSON as I want to make it as easy as possible for people to create new workflows. Plus, they also get versioned easily this way.
simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
Sure!

In Agentflow, you write functions by inheriting from the BaseFunction class. You need to provide the definition in JSON that GPT-3.5/4 uses to understand how to call a function, and also the function logic itself. This just means creating a get_definition() function that returns a JSON Schema object, and an execute() function that performs your logic and returns a string. Once you have those, you can then just use the function in your workflow by adding "function_call": "your_function". The application does the rest. Here's the create_image function, for example, which uses the Dall-e API: https://github.com/simonmesmith/agentflow/blob/main/agentflo...

What I mean by "you don't need to write any code with LangChain" is that you don't need to write any Python at all to use Agentflow, unless you want to create a new function. Creating workflows just involves creating JSON files. It's not like LangChain, for which you'd have to chain together multiple prompts in Python.

Does that help clarify?

PS: You'll notice heavy documentation in the link above. I want to experiment with automatically generating documentation using Sphinx, so I documented everything with Sphinx formatting. It might be overkill.

simonmesmith··on Show HN: Agentflow – Run Complex LLM Workflows from Simple JSON
For sure.

Let's use an example, like this: https://github.com/simonmesmith/agentflow/blob/main/agentflo...

This is a workflow for coming up with a product idea and illustrating it with an image.

This workflow has 9 steps, starting with brainstorming ideas, and ending with saving an HTML file containing the product name, description, and image.

It also has two variables, {market} for the target market, and {price_point} for the price point.

To run this workflow, you simply enter this in the command line: python -m run --flow=example_with_variables --variables 'market=college students' 'price_point=$50'

You don't need to write any code with LangChain.

You simply specify your workflow in a JSON file, and execute it.

Does that help to clarify?

simonmesmith··on Exploring the effects of feeding emotional stimuli to large language models
Regarding your last comment, I think it rings true, and further, politics (or at least political communication) is to a degree citizen prompt engineering.

It further strikes me that, if these are true:

(1) Influencing large groups of people is at least somewhat equivalent to prompt engineering. (2) Large language models reflect at least some degree of similarity with how humans respond to the same prompts.

Then you could use LLMs to test the influence of different psychological manipulations before using them on large groups of people.

Try different prompts, see which one gets the best response from an LLM, then use it on people.

simonmesmith··on Working Hours Data
It would be interesting to see this chart start much further back, such as when most humans were hunter-gatherers. Yes, working hours per week have declined since the beginning of agriculture, but there is evidence that humans worked much more as farmers than as hunter-gatherers. For example: https://www.cam.ac.uk/research/news/farmers-have-less-leisur....
simonmesmith··on Make way for Solein, a new vegan protein on the menu
A bit of sleuthing from the company’s website to Wikipedia to Google Patents brought me to this: https://patents.google.com/patent/WO2021084159A1/en

It is not peas. It is a specific bacteria. They just talk about a microorganism on the website, probably because people don’t want to think about eating bacteria (though they eat yogurt just fine).

simonmesmith··on Ask HN: Is it too soon to rely on LLM frameworks? Better to code it myself?
When I first started experimenting with LLMs, I used LangChain. It helped me get started, but over time I found that:

(1) Things were often more complicated than they had to be. For example, creating prompt template objects rather than just using a simple format call to replace variables in a string.

(2) Too much of the underlying mechanics got hidden away. For example, creating agents hides away the prompting to get them to use functions, so it’s hard to know why things are going wrong.

(3) Different LLMs aren’t interchangeable, so adding an abstraction for the purpose of enabling plug-and-play with different providers doesn’t really help. For example, OpenAI has function calls and streaming output, while Anthropic has a 100K context window. Those differences make abstraction less compelling, because not all LLMs are interchangeable.

(4) It limited my creativity. I found I was trying to make my code conform to LangChain’s implementations.

(5) It made things heavier than they needed to be. Instead of just installing, say, the OpenAI library, I would install LangChain with all of its dependencies.

(6) It made it harder for people to see how my code was working. They had to understand how LangChain was doing what it was doing, in addition to what my code was doing.

(7) The OpenAI API (which I use primarily) is already pretty easy to work with. Other LLM-related libraries, like Chroma, or Pinecone, are also easy to work with. So why do I need LangChain?

My two cents, for what it’s worth.

simonmesmith··on G/O Media Incorporating AI-Generated Content
I agree.

I also keep thinking: If ChatGPT (etc.) can write such content, and I'm interested in it, then why won't I just ask ChatGPT for the information directly rather than going to someone's bloated website?

simonmesmith··on Instant Coffee Is Negatively Associated with Telomere Length
Based on UK Biobank data and self-reports. Possible reason is: “The mineral lead in instant coffee was more abundant than that in other coffee types, and long-term consumption of instant coffee may result in excessive lead. Additional substances added to commercial instant coffee, such as creamer and flavoring agents, might partially explain the negative effect.” Additives aren’t accounted for here.
simonmesmith··on Vegetarian or vegan diets and blood lipids: a meta-analysis of randomized trials
TL; DR: “Compared with the omnivorous group, the plant-based diets reduced total cholesterol, low-density lipoprotein cholesterol, and apolipoprotein B levels with mean differences of −0.34 mmol/L (95% confidence interval, −0.44, −0.23; P = 1 × 10−9), −0.30 mmol/L (−0.40, −0.19; P = 4 × 10−8), and −12.92 mg/dL (−22.63, −3.20; P = 0.01), respectively. The effect sizes were similar across age, continent, duration of study, health status, intervention diet, intervention program, and study design. No significant difference was observed for triglyceride levels.”
simonmesmith··on Ask HN: Advice from people who strength train from home
For what it’s worth, I’ll mention what works for me. I have no interest in any companies or products mentioned below other than using them and finding them useful.

I’ve weight-trained for decades and switched up my routine during the pandemic. I have only a small room available at home for this, which I also use as an office and music studio. So, not a lot of space.

I bought a pair of Bowflex SelectTech 552s (https://www.bowflex.com/selecttech/552/100131.html), a stand for them, a bench, a stability ball, and mini loop bands. Then I started using the Fitbod app (https://fitbod.me) to take the guesswork out of my routine.

This has been great. Fitbod will gradually ramp up resistance, vary your workouts, target different muscles based on recovery, and adjust your weight levels based on your performance. Then you can use the 552s to select the right weight it recommends.

Anyway, your mileage may vary, but I’ve been very happy with this setup for years.

simonmesmith··on AITodo: Todo app powered by a GPT backend
This is a really simple but elegant idea, and from a few other examples I’ve seen it feels very much like there are lots of fun ideas to explore in the space of flexible GPT-based APIs.

Your idea, for example, prompted an idea for a GPT-based recipe app. You just say something like, “provide a recipe for macaroni and cheese and add the ingredients to my shopping list.” Then you get a shopping list using your to-do approach.

simonmesmith··on Ask HN: Have you made a career move “down” on purpose, and how has it been?
I was a communications executive and had been for more than a decade. But I always had a passion for engineering and built many side projects. I decided to pursue that passion and became an entry-level engineer. I was fortunate to have my company’s (and my wife’s) support. I went from managing a team of about 20 to being part of a small R&D team of about 6. It was the right change for where I’m at in life now despite a substantial cut in compensation.
simonmesmith··on Maximum two drinks a week, Canada advises
As a Canadian, I think it’s important to provide some additional context. In Ontario, the province where I live, the provincial government has a near monopoly on alcohol sales through its LCBO stores. These stores bring in well over $2 billion of revenue each year (https://www.lcbo.com/content/dam/lcbo/PDFs/AnnualReport/LCBO...). While I know there are arguments for this like “if people are going to drink anyway, the government should control and benefit from it,” it certainly sends mixed messages. The LCBO even advertises to promote alcohol consumption. Even worse, the sale of alcohol doesn’t offset the costs of alcohol to governments, according to this study on a government website: https://www.canada.ca/en/public-health/services/reports-publ...
simonmesmith··on GPTZero
Maybe I’m naive here, but whenever I hear about a new AI fake detector, I think about how I would use it in some kind of generative adversarial network or reinforcement learning setup whereby a neural network improves at generating non-AI-sounding material based on feedback from these detectors. It just feels like an arms race, and the existence of these fake detectors will simply help generative AI improve faster. No?
simonmesmith··on CNET Money: financial 'advice', AI generated, fact-checked by editorial staff
Wow. There were 75 articles on the page when I read it. As a former journalist, I would be curious to know what “fact-checked” actually means. I would also be curious to compare these “fact-checked” articles to direct output from ChatGPT. If the latter is already producing content that’s close to equal in factual accuracy, then why would anyone read these articles rather than just have a conversation with ChatGPT?
simonmesmith··on Ask HN: Are you also using ChatGPT a lot every day for code related work?
I’ve been experimenting with both ChatGPT and Copilot. I find that Copilot is much more accurate and context-aware, but I can’t use it everywhere—e.g. when working in a shared Colab notebook. Conversely, I find that ChatGPT is better if I’m brainstorming some ideas and want to get a high-level sense of different ways to execute them in code.

This said, ChatGPT has a tendency to get things about 95% right, so you have to be very careful. For example, I was playing with it yesterday and it wrote a Python function with regex that didn’t do what I had asked, even though it said it did. I pointed out the error and it rewrote the regex but made it worse.

Oddly, when I started another thread and asked ChatGPT just for the regex outside of a Python function, it nailed it. So I do think prompt engineering can overcome some of the limitations I’ve seen to date.

simonmesmith··on Ask HN: Is it normal to like people at a distance and dislike otherwise?
I can relate to this. For me, I think it comes from a combination of wanting meaningful conversation (which doesn’t need to happen too frequently because you run out of meaningful things to discuss), having minimal interest in networking based on where I’m at in my career, having less time for frivolous meetups because I have many work and family obligations and want more time to myself, appreciating other things like spending time in nature and wanting to devote more time there, and being more discerning about who I spend time with.

It may change again, and perhaps there will be a point where I want to see people more frequently and frivolously. But right now, I have a strong urge similar to what you describe.

simonmesmith··on ChatGPT Tackles xkcd's 'What If'?
Having recently seen a paper on the benefit of chain-of-thought prompting, I wondered how this might differ if you added something like “thinking through this step by step…” to the prompts.

I tried it with the mountain prompt, as follows:

> Thinking through this step by step: What if a huge mountain—Denali, say—had the bottom inch of its base disappear? What would happen from the impact of the mountain falling 1 inch? What about 1 foot? What if the mountain's base were raised to the present height of the summit, and then the whole thing were allowed to drop to the earth?

ChatGPT’s response:

> It's not possible for the bottom inch of the base of a mountain like Denali to suddenly disappear, as the mountain is connected to the Earth's crust and is held in place by the planet's gravity. However, if the entire mountain were somehow lifted up and then allowed to drop, the impact would be catastrophic. The energy released by such an event would be immense, and it would likely cause widespread damage and destruction. It's also worth noting that the Earth's crust is not solid and uniform, so the ground beneath the mountain would likely be deformed by the impact.

I think we’re learning just how sensitive large language models are to prompting.

simonmesmith··on Can You Recommend a Toaster?
This isn’t exactly what you’re looking for, but a few months ago we had to buy a new toaster oven and I went online to find reviews. They pointed me, surprisingly, to an infrared toaster oven by Panasonic (this one: https://shop.panasonic.com/kitchen-and-home/kitchen-applianc...). It is phenomenal. Fast, even toasting and even baking. The controls kinda suck—they feel like they’re from the 1970s—but it works incredibly well.
simonmesmith··on Social media may prevent users from reaping creative rewards of profound boredom
This looked interesting and then I read the following in the press release: “Dr Hill said the research sampled 15 participants of varying age, occupational and education backgrounds in England and the Republic of Ireland, who had been put on furlough or asked to work from home.”

So the hypothesis is based on 15 people, in one region of the world, in a very specific circumstance. The authors admit the research is limited, but if you don’t read deeply into the press release you might come away thinking this finding—if you can call it that—is much more solid than the actual survey methodology would support.

simonmesmith··on Study uncovers factors linked to optimal aging
What are these magical factors?

“They found that those who were female, married, physically active and not obese and those who had never smoked, had higher incomes, and who did not have insomnia, heart disease or arthritis, were more likely to maintain excellent health across the study period and less likely to develop disabling cognitive, physical, or emotional problems.”

So, basically, be a woman, get married, be rich, eat well, exercise, don’t smoke, and, oh yeah, don’t be sick to begin with:

“As a baseline, the researchers selected participants who were in excellent health at the start of the approximately three-year period of study. This included the absence of memory problems or chronic disabling pain, freedom from any serious mental illness and absence of physical disabilities that limit daily activities—as well as the presence of adequate social support and high levels of happiness and life satisfaction.”

But wait, is that even the actual journal article’s focus? Umm, it seems to focus on immigrant Canadians, with the conclusion: “Immigrant older adults had a lower prevalence of successful aging than their Canadian-born peers.”

So, my cynical take is:

1. Big study funded 2. Needs to produce some kind of compelling results for the public 3. Actual conclusions aren’t particularly insightful 4. PR produced in attempt to spin the results

simonmesmith··on Toward customizable timber, grown in a lab
Maybe this is an interesting scientific development, but to pitch it as aiding deforestation is a massive stretch. Three quarters of deforestation is due to agriculture: https://ourworldindata.org/what-are-drivers-deforestation. So cultured meats would probably have a much bigger impact on deforestation than cultured trees.
simonmesmith··on Ask HN: Why isn't remote work advertised as a pro environment initiative?
I don’t know about universally true, so you may be right that there are exceptions, but as one example, such as linked above, public transit tends to underserve the needs of female caregivers who often need to make multiple short trips rather than two long trips each day.
simonmesmith··on Ask HN: Why isn't remote work advertised as a pro environment initiative?
Similarly, I often wonder why cities don’t promote local telecommuting over building expensive public transportation infrastructure. Building one kilometre of subway costs anywhere from around $100 million to $1.6 billion (https://marroninstitute.nyu.edu/events/how-much-does-it-cost...), and then requires subsidies to run. Why not use this money to provide incentives for people to telecommute?

And while some might say, “but then people who rely on public transit for things other than commuting to work would be hurt,” note that most public transit currently doesn’t serve those people well because it’s focused on work-related commuting patterns. Example: https://www.ualberta.ca/folio/2022/01/public-transit-service....

← PreviousPage 2 of 3Next →