185 karma · joined August 13, 2022
The main issue right now is that I’m relying on the OpenAI API’s function-calling capabilities to enable the use of functions in workflows.
If we switch to other LLMs, we’ll need to create some additional wrapping around them to allow such function calling. (As far as I know. If any open source LLMs already have function calling built in, let me know.)
LangChain (which I’m not using due to personally finding it overcomplicated and too obfuscating) does have function calling, but it uses approaches like REACT (if I remember correctly) that aren’t as reliable as OpenAI’s approach.
But I think there are some key differences. I’m not sure if Magic Loop is open source, for example, so I don’t know how they’re currently building their workflows.
Automating API calls is neat, too, but I get a bit anxious with too much automation when you want to run a workflow with consistent results. My gut says you want to have the functions here pretty locked down and tested, rather than rely on automating API calls. But maybe I’ll be proven wrong.
In Agentflow, you write functions by inheriting from the BaseFunction class. You need to provide the definition in JSON that GPT-3.5/4 uses to understand how to call a function, and also the function logic itself. This just means creating a get_definition() function that returns a JSON Schema object, and an execute() function that performs your logic and returns a string. Once you have those, you can then just use the function in your workflow by adding "function_call": "your_function". The application does the rest. Here's the create_image function, for example, which uses the Dall-e API: https://github.com/simonmesmith/agentflow/blob/main/agentflo...
What I mean by "you don't need to write any code with LangChain" is that you don't need to write any Python at all to use Agentflow, unless you want to create a new function. Creating workflows just involves creating JSON files. It's not like LangChain, for which you'd have to chain together multiple prompts in Python.
Does that help clarify?
PS: You'll notice heavy documentation in the link above. I want to experiment with automatically generating documentation using Sphinx, so I documented everything with Sphinx formatting. It might be overkill.
Let's use an example, like this: https://github.com/simonmesmith/agentflow/blob/main/agentflo...
This is a workflow for coming up with a product idea and illustrating it with an image.
This workflow has 9 steps, starting with brainstorming ideas, and ending with saving an HTML file containing the product name, description, and image.
It also has two variables, {market} for the target market, and {price_point} for the price point.
To run this workflow, you simply enter this in the command line: python -m run --flow=example_with_variables --variables 'market=college students' 'price_point=$50'
You don't need to write any code with LangChain.
You simply specify your workflow in a JSON file, and execute it.
Does that help to clarify?
It further strikes me that, if these are true:
(1) Influencing large groups of people is at least somewhat equivalent to prompt engineering. (2) Large language models reflect at least some degree of similarity with how humans respond to the same prompts.
Then you could use LLMs to test the influence of different psychological manipulations before using them on large groups of people.
Try different prompts, see which one gets the best response from an LLM, then use it on people.
It is not peas. It is a specific bacteria. They just talk about a microorganism on the website, probably because people don’t want to think about eating bacteria (though they eat yogurt just fine).
(1) Things were often more complicated than they had to be. For example, creating prompt template objects rather than just using a simple format call to replace variables in a string.
(2) Too much of the underlying mechanics got hidden away. For example, creating agents hides away the prompting to get them to use functions, so it’s hard to know why things are going wrong.
(3) Different LLMs aren’t interchangeable, so adding an abstraction for the purpose of enabling plug-and-play with different providers doesn’t really help. For example, OpenAI has function calls and streaming output, while Anthropic has a 100K context window. Those differences make abstraction less compelling, because not all LLMs are interchangeable.
(4) It limited my creativity. I found I was trying to make my code conform to LangChain’s implementations.
(5) It made things heavier than they needed to be. Instead of just installing, say, the OpenAI library, I would install LangChain with all of its dependencies.
(6) It made it harder for people to see how my code was working. They had to understand how LangChain was doing what it was doing, in addition to what my code was doing.
(7) The OpenAI API (which I use primarily) is already pretty easy to work with. Other LLM-related libraries, like Chroma, or Pinecone, are also easy to work with. So why do I need LangChain?
My two cents, for what it’s worth.
I also keep thinking: If ChatGPT (etc.) can write such content, and I'm interested in it, then why won't I just ask ChatGPT for the information directly rather than going to someone's bloated website?
I’ve weight-trained for decades and switched up my routine during the pandemic. I have only a small room available at home for this, which I also use as an office and music studio. So, not a lot of space.
I bought a pair of Bowflex SelectTech 552s (https://www.bowflex.com/selecttech/552/100131.html), a stand for them, a bench, a stability ball, and mini loop bands. Then I started using the Fitbod app (https://fitbod.me) to take the guesswork out of my routine.
This has been great. Fitbod will gradually ramp up resistance, vary your workouts, target different muscles based on recovery, and adjust your weight levels based on your performance. Then you can use the 552s to select the right weight it recommends.
Anyway, your mileage may vary, but I’ve been very happy with this setup for years.
Your idea, for example, prompted an idea for a GPT-based recipe app. You just say something like, “provide a recipe for macaroni and cheese and add the ingredients to my shopping list.” Then you get a shopping list using your to-do approach.
This said, ChatGPT has a tendency to get things about 95% right, so you have to be very careful. For example, I was playing with it yesterday and it wrote a Python function with regex that didn’t do what I had asked, even though it said it did. I pointed out the error and it rewrote the regex but made it worse.
Oddly, when I started another thread and asked ChatGPT just for the regex outside of a Python function, it nailed it. So I do think prompt engineering can overcome some of the limitations I’ve seen to date.
It may change again, and perhaps there will be a point where I want to see people more frequently and frivolously. But right now, I have a strong urge similar to what you describe.
I tried it with the mountain prompt, as follows:
> Thinking through this step by step: What if a huge mountain—Denali, say—had the bottom inch of its base disappear? What would happen from the impact of the mountain falling 1 inch? What about 1 foot? What if the mountain's base were raised to the present height of the summit, and then the whole thing were allowed to drop to the earth?
ChatGPT’s response:
> It's not possible for the bottom inch of the base of a mountain like Denali to suddenly disappear, as the mountain is connected to the Earth's crust and is held in place by the planet's gravity. However, if the entire mountain were somehow lifted up and then allowed to drop, the impact would be catastrophic. The energy released by such an event would be immense, and it would likely cause widespread damage and destruction. It's also worth noting that the Earth's crust is not solid and uniform, so the ground beneath the mountain would likely be deformed by the impact.
I think we’re learning just how sensitive large language models are to prompting.
So the hypothesis is based on 15 people, in one region of the world, in a very specific circumstance. The authors admit the research is limited, but if you don’t read deeply into the press release you might come away thinking this finding—if you can call it that—is much more solid than the actual survey methodology would support.
“They found that those who were female, married, physically active and not obese and those who had never smoked, had higher incomes, and who did not have insomnia, heart disease or arthritis, were more likely to maintain excellent health across the study period and less likely to develop disabling cognitive, physical, or emotional problems.”
So, basically, be a woman, get married, be rich, eat well, exercise, don’t smoke, and, oh yeah, don’t be sick to begin with:
“As a baseline, the researchers selected participants who were in excellent health at the start of the approximately three-year period of study. This included the absence of memory problems or chronic disabling pain, freedom from any serious mental illness and absence of physical disabilities that limit daily activities—as well as the presence of adequate social support and high levels of happiness and life satisfaction.”
But wait, is that even the actual journal article’s focus? Umm, it seems to focus on immigrant Canadians, with the conclusion: “Immigrant older adults had a lower prevalence of successful aging than their Canadian-born peers.”
So, my cynical take is:
1. Big study funded 2. Needs to produce some kind of compelling results for the public 3. Actual conclusions aren’t particularly insightful 4. PR produced in attempt to spin the results
And while some might say, “but then people who rely on public transit for things other than commuting to work would be hurt,” note that most public transit currently doesn’t serve those people well because it’s focused on work-related commuting patterns. Example: https://www.ualberta.ca/folio/2022/01/public-transit-service....