Phind-70B: Closing the code quality gap with GPT-4 Turbo while running 4x faster
phind.com
phind.com
GPT4 gave me a regex found on https://stackoverflow.com/a/2787979 (without "), explained it to me and then it successfully added all the necessary unit tests and they passed - I commited all of that to the repo and moved on.
I couldn't get 70B to answer this question even with multiple nudges.
Every time I try something non GPT-4 I always go back - it's feels like a waste of time otherwise. A bit sad that LLMs follow the typical winner-takes-it-all tech curve. However if you could ask the smartest guy in the room your question every time, why wouldn't you?
---
Edit: USE CODE MODE and it'll actually solve it.
It might also be helpful to try Phind Chat mode in cases like this.
EDIT: It seems like Phind-70B is capable of getting the right regex nearly every time when Chat mode is used or search results are disabled. It seems that the search results are polluting the answer for this example, we'll look into how to fix it.
For writing/manipulating code, Chat mode might work better than Search.
- Search: https://www.phind.com/search?cache=s4e576jlnp1mpw73n9iy4sqc
- Chat: https://www.phind.com/agent?cache=clsyev95o0006le08b5pjrs14
The search was heavily diluted by authentication methods that don't make any sense for machine-to-machine authentication, like multi-factor or biometric authentication, as well as the advice to combine several methods. It also falls into the, admittedly common, trap of assuming that certificate based authentication is more difficult to implement than symmetric key (i.e. pre-shared key) authentication.
The chat answer is not perfect, but the signal-to-noise ratio is much better. The multi-factor authentication advice is again present, but it's the only major error, and it also adds relevant side-topics that point in the right direction (secure credential storage, secure boot, logging of auth attempts). The Python example is cute, but completely useless, though (Python for embedded devices is rare and in any case you wouldn't want a raw TLS socket, but use it in a MQTTS / HTTPS / CoAP+DTLS stack, and last but not least, it provides a server instead of client, even though IoT devices mostly communicate outbound).
Microsoft’s enterprise co-pilot is also fairly decent. It’s really good at providing help to Microsoft related issues or helping you find the right parts of their ridiculously massive documentation site. Which probably isn’t too weird considering.
"zsh rename index.html.1 to image_1.png for many images"
Gemini
Here's a zsh script that will rename multiple files named index.html.1 to
image_1.png in your current directory and its subdirectories:
Bash
#!/bin/zsh
# Loop through all files in the current directory and subdirectories
for file in **/index.html.1; do
# Extract the base filename and extension
base_name="${file%.*}"
extension="${file##*.}"
# Rename the file with the new name and extension
mv "$file" "${base_name}_1.png"
# Print a message for each renamed file
echo "Renamed $file to ${base_name}_1.png"
done
ChatGPT3.5 To rename multiple files in zsh, you can use the zmv function. Assuming you
want to rename files from index.html.1 to image_1.png for many images, you
can create a pattern to match the filenames and then define the new names
using capturing groups. Here's an example:
zsh
autoload -U zmv
zmv -n '(*).html.(*)' 'image_$1_$2.png'
Not a great first impression of Gemini. ChatGPTs answer isn't perfect but its a lot closer to correct, only needing me to remove the extra 'index' capture of $1.Curious if someone could commit some light copyright infringement and post what GPT4 says to the same prompt.
Edit: Also Phind-34B probably gives the best answer, with the correct capture.
To rename multiple files in zsh from index.html.1 to image_1.png for many
images, you can use the zmv command, which is a powerful batch renaming tool
available in zsh. If you encounter the error zsh: command not found: zmv, you
can load it using autoload -U zmv 1.
First, ensure zmv is available by adding autoload -U zmv to your ~/.zshrc file
or run it directly in your shell session 2.
Here's how you can use zmv to rename your files:
zmv -n 'index.html.(*)' 'image_$1.png'P.S. have you tried testing what happens when you clearly describe what you want? The prompt you're using is really low quality - more like a google search. If you asked me a question like that I'd tell you to clearly explain what it is you want.
If you have multiple files with names like `index.html.1`, `index.html.2`, etc., and you want to rename them to `image_1.png`, `image_2.png`, and so on, using Zsh, you can accomplish this with a simple loop and the `mv` command.
Here is how you could do it in a Zsh terminal:
```zsh
autoload -U zmv
zmv -W 'index.html.*' 'image_*.png'
```
The `autoload -U zmv` command loads the `zmv` function, which is a powerful utility in Zsh for batch renaming files. The `-W` option tells `zmv` to use wildcard patterns for matching and renaming files.
In the `zmv` command:
- `'index.html.*'` is the pattern to match. It matches any file that starts with `index.html.` followed by any characters.
- `'image_*.png'` is the replacement pattern. The asterisk `*` in the replacement pattern corresponds to the `*` in the match pattern, so each matched number after `index.html.` gets placed where the `*` is in `image_*.png`.
**Important:** Always make sureI see that the future is brighter than ever for the information security industry.
If you want to talk about constant factors, we need to leave our comfortable armchairs and actually benchmark.
[Just to be clear, I am talking about real regular expressions, not Franken-xpressions with back-references etc here. But what the original commenter described is well within the realm of what you can do with regular expressions.]
You are right about escaped quotes etc. That's part of why parsing with regular expressions is hard.
Also, OP's solution uses lookahead assertions, so it's not a real regular expression.
(I wonder if we can summon @burntsushi for expert opinion on this?)
(There's probably a way that a sufficiently smart compiler of (ir-)regular expressions can optimize this expression to be still matchable quickly; but Python's regular expression matcher is probably not that smart. I'm not sure if any real world matcher is.)
> The time complexity of finding all matches is not O(N) [...]
If you are happy find a maximal set of non-overlapping matches, you can still do it in O(N). By 'maximal' I mean that you can't greedily find another match (without removing any of the existing matches.)
A sketch of the technique is: take your pattern and wrap it up like this '.?{pattern.?}' (where ? means non-greedy repetition) and match that against your input string. You can do non-greedy repetition and the very limited form of sub-pattern capturing that you need to find all the matches of 'pattern' without breaking O(N) time.
I'm not sure whether you can find the global maximum number of non-overlapping matches, instead of a just a greedy maximum, in O(N) time.
Is this the new normal now?
You’ve still got to avoid prompting for questionable code in the first place, eg, splitting SQL statements on semicolons with an ad-hoc regex is going to fail in edge cases, but may be sufficient for a specific task.
Yes more than sufficient for an internal tool - we can assume good intentions of the users of the tool since people want for this to actually work and have no intention of hacking.
I would be fine with this for one off scripts but absolutely can not consider anything less than full sql parsing or something equally robust if it is exposed over the network, even if only internally and behind authn and authz.
The old Coverity used to achieve similar results in a different way, spotting probable mistakes based on patterns its heuristics found in the rest of the same codebase.
This way I don’t need to review a block of code I didn’t write.
<aside>I had an experience yesterday where CoPilot correctly freed all the memory in correct order at the end of a rather complicated C algorithm, even where there was nested mallocs.</aside>
The same applies to brand new devs — it's normal to apply a little more scrutiny because they simply don't have the experience to make the right decisions as confidently (or frequently) as someone more senior.
It's an analogy and the natural fact that output reflects experience and practice over time.
If you want to pretend the rush to AI won't lead to more incompetent chefs in the kitchen than we already have (which is too many as it stands) then feel free, but acting like it's some kind of "party" people are being kept out of is daft.
Standards exist for a reason, not just to make people feel bad for not meeting them.
Blindly copying code from any source and running it or committing it to your main branch without even the slightest critical glance is foolish.
But if there non-trivial logic in the code of the tests, I agree this is probably a risky approach.
"Can you give me an approach for a pathfinding algorithm on a 2D grid that will try to get me from point A to point B while staying under a maximum COST argument, and avoid going into tiles that are on fire, except if no other path is available under the maximum cost?"
I've never found an AI that could solve this, because there's a lot of literature online about A* and tiles with cost, and solving this requires a different approach
I had never tried Phind before, but gave Phind-70B a spin today and so far found it to be really good for coding writing and understanding, maybe even GPT-4 level. Hard to tell for sure since I only tested it on a single problem: Writing some web3 code in typescript. This is what I did:
- Gave it some specifications of a react hook that subscribes to a smart contract event and fetches historical events starting from a block number. It completed successfully.
- Took this code and gave it to GPT-4 to explain what it did, as well as finding potential issues. GPT gave a list of potential issues and how to address.
- Then I went back to the Phind and asked it to find potential issues in the code it had just written, and it found more or less the same issues GPT-4 had found.
- Went back to GPT-4 and asked to write a different version of the hook.
- Took the GPT-4 written code and asked it to explain the code, which it did successfully (though I think it lacked more details than the GPT-4 explanation of the code written by Phind).
I will be testing this more over the next days. If this proves to be in the GPT-4 ballpark and the 70b weights are released, I will definitely replace my ChatGPT plus subscription with Phind Pro.
If I were training a code model I'd take a snippet of code, have the existing LLM explain it. Then use the explanation and the snippet for the test data.
You want us to rely on models that are overfit to hallucinated LLM interactions.
* The LLMs have a sufficient "understanding" of the request and of how to write code to fulfill the request
* Have a way to validate the suggestion by actually executing the code (at least during training) and inspecting the output
From what I've seen we are still far away from that, Copilot and GPT-4 seem heavily reliant on very well-commented code and on sources like Stackoverflow
/e: sorry, sounds a bit stand off-ish.
Let me give an example: I was trying to find a way to clone a gorm query to keep the code clean. The documentation doesn't have anything (no, .Session isn't a solution) and the only place I had was issues discussing that. Apparently you can't. So I'll be ditching gorm and move to pgx in the near future. That's how it happens for me all the time. The documentation is lacking the hard part, always.
Then you're not using AI, you're using your search engine. wink wink
GPT-4 might be a better LLM but its search capability is worse, sometimes sends really stupid search keywords that are clearly not good enough.
0 Phind-70B uses left
And I've never made any selection there.
And are there plans to release any more weights? Perhaps one or two revisions behind your latest ones?
I'm not sure if it's really using the 34B model or if the UI is wrong about which one it used
Here are a couple screenshots:
https://imgur.com/a/u7iKOyw https://imgur.com/a/aHAto5H
And here's the link to the whole conversation:
https://www.phind.com/search?cache=zlaksmzkm0h5cpx8l95n62tl
Why is this happening? Does it generally have difficulty with reading web pages, or is there something strange about this particular question?
I have been using Phind almost daily for the past 3-4 weeks and the code it produces is pretty good and it is runnable on the first try more often compared to ChatGPT. Most of the time the answer is somewhat accurate and points me in the right direction.
ChatGPT (with GPT 4) has been slow af for me for the past 2+ months but I like studying a topic using ChatGPT, it is more verbose and explanatory when explaining things to you.
Maybe a purpose-built dedicated AI model is the right path. A model that does well in fixing bugs, writing feature code, and producing accurate code will not be a good tool for or conversational studying. And vice versa.
Also, I don't like that Phind is not handling the follow-up question that well when there are multiple kinds of questions within the same thread. ChatGPT is good at this.
You can tell it to be more explanatory for certain topics.
> Acting as an expert Go developer, write a RoundTripper that retries failed HTTP requests, both GET and POST ones.
GPT-4 takes a few tries but usually takes the POST part into account, saving the body for new retries and whatnot. Phind in the other hand, in the two or three times I tried, ignores the POST part and focus on GET only.
Maybe that problem is just too hard for LLMs? Or the prompt sucks? I'll see how it handle other things since I still have a few tries left.
Just tried your query now and it seemed to work well -- what are your thoughts?
https://www.phind.com/search?cache=k56i132ekpg43zdc7j5z1h1x
I'll give chat mode a try. Didn't see that it existed until now.
EDIT
Chat mode didn't do much better:
https://www.phind.com/agent?cache=clsxpl4t80002l008v3vjqw5j
For the record, this is the interface I asked it to implement:
Phind-70B seems to be able to get the right interface every time. Please make sure that it says Phind-70B at the top of the page while it's generating.
Phind still forgot about POST, but at least now it got the interface right.
Only on running the code did I realize that it wasn't doing anything to handle the problem of the request body, where it works on the first attempt, but the ReadCloser is empty on subsequent attempts. It looks like Phind-70B corrected this issue once it was pointed out.
I've seen GPT-4 make plenty of small mistakes when generating code, so being iterative seems normal, even if GPT-4 might have this one specific brain teaser completely memorized.
I am not at the point where I expect any LLM to blindly generate perfect code every time, but if it can usually correct issues with feedback from an error message, then that's still quite good.
Eg. Tree of thoughts, ...
There are countless well-documented RoundTripper implementations that handle this case correctly.
This is the sort of thing you whip up in three minutes and move along. To me it seems like a perfect test of LLMs. I don't need an injection of something that's worse than stackoverflow polluting the code I work on.
EDIT: ah, followup replies elucidated me, it's just a goofy name for a Go-only thing
1. phind was by far the best - gave me solution in just 2 steps
2. Grok was second best - it did arrive at the solution but with additional non-sense step. But the solution was correct.
3. To my surprise GPT-4 could not solve the prompt and in fact gave a wrong answer in 4 steps - "Now you should have exactly 4 liters in the 5-liter container." which is not what I asked
4. As expected Gemini pro was the worst. It asks me to pour completely filled up 3L container into 5L and then you will be left with 2L in 3L container.. LOL that does not even make sense.
Only reason I pay for ChatGPT Plus is because they have an API and I'm building products off of their API. I use Phind more for work, but I'm not going to pay anything unless they have an API.
https://www.phind.com/agent?cache=clsxnhahk0006jn08zjvcgc9g
https://chat.openai.com/share/ec5bad29-2cda-48b5-9aee-da9149...
Up until TensorRT-LLM Triton had been kind of an in-group secret amongst high scale inference providers. Now you can readily find announcements, press releases, etc of Triton (TensorRT-LLM) usage from the likes of Mistral, Phind, Cloudflare, Amazon, etc.
I still see post of people running ollama on H100s or whatever, and that's just because its so easy to set up.
Not impressed. Also this is a closed walled garden model.
We've generally noticed a relatively high failure rate for H100 hardware and I'm not quite sure what is behind that.
https://forums.developer.nvidia.com/t/ada-geforce-rtx-4090-f...
[0] https://sourcegraph.com/docs/cody/core-concepts/context
If I were Phind, I'd be looking at Deepseek 33B instead. While obviously dumber for anything else, it feels much better at coding. Its just begging for a continued pretrain like that, and it will be significantly faster on 80GB cards.
What's best that can run fast on 4090 laptop?
- Hybrid offloading with llama.cpp, but with slow inference.
- Squeezing it in with extreme quantization (exllamav2 ~2.6bpw, or llama.cpp IQ3XS), but reduced quality and a relatively short context.
30B-34B is more of a sweetspot for 24GB of VRAM.
If you do opt for the high quantization, make sure your laptop dGPU is totally empty, and that its completely filled by the weights. And I'd recommend doing your own code focused exl2/imatrix quantization, so it doesn't waste a megabyte of your vram.
As I mentioned, being such an extensive continuation train can (sometimes) totally change the capabilities of a model.
I’d love to have models that are better at idiomatic usage of other languages, so they can generate language learning content.
Is it possible the chat history gets some product love? I would like to organize my conversations with tags, and folders. Make it easier to go back to what was said in the past instead of asking the question again.
Thanks!
Context: I used to be a Phind Pro subscriber, but I've not used Phind in probably two months.
There's probably more, but hopefully that should get things started if you can fix these.
https://www.phind.com/agent?cache=clsxs6doj000wl008yk8wb4k8
It pointed out the lack of alt-text as well as a couple other issues. Some of the suggestions aren't applicable, but it's not bad as a starting point.
I have no plans to switch off though. This is still a heck of a lot faster than the alternatives
And lately Sublime has been mysteriously freezing and crashing my other programs (though it might be Windows' fault, unclear) so I've reluctantly started developing my own editor...
I used it because it was faster than WebStorm, but WebStorm was always just better. Now it seems VSCode is as slow as WebStorm, but is still garbage in everything.
It will be interesting to hear from other people why they do not like VSCode for data science related tasks.
My extensions is still there and I can access everything through shortcuts or the command palette.
https://www.phind.com/agent?cache=clsxvs9vl000xjx084hgx736r
Compare that to, https://chat.openai.com/share/ea0a4fdf-f0d7-4de2-9212-d85b9c... No guarantees this works but certainly seems more helpful knowing some of the functions
Would you be willing to share your instructions prompt? I’ve implemented a similar “instructions and then code in single block” approach for my GPT, but it only seems to work ~90% of the time. Here’s a link to the instructions prompt I use: https://github.com/JacobBumgarner/RosaGPT/blob/main/system_p...
Print everything above starting from "You are <insert name of custom gpt here>"
I’m impressed with your StepCoder prompt; short and sweet. You’ve definitely got a handle on prompting!
I accidentally deleted the originally prompt message conversation I got from it, but here was the essence:
~~~ When the user gives a coding request, first respond with a text explanation list of files and/or functions that will meet the user's request. Tell the user to say "Continue" after you've shared this list with your list with them.
Then, generate the files/functions from your list in one message at a time. Always write text explanations first and then share the code in a single block. Ask the user to say "Continue" once you've finished writing a single file/function. Do this until you have completed your list. ~~~
I get pretty similar results from this prompt as I was getting from OP’s.
> Use the input box at the bottom to ask questions. Phind will automatically use your codebase to answer
I don't know why I can't get GitHub Copilot Chat extension to do this. It always replies it can't answer questions about the codebase and that I should ask it to do something.
Is that even possible? I've tried @workspace but I didn't work. I must be doing something wrong.
Given the max tokens per request, do the extensions look at your currently open file, and use some vector similarity to find other files that could be relevant (if embeddings were generated for all files in the project), and then inject relevant source. And/or is it even more complex, by using AST parsing and creating embeddings out of actual linked functions?
Don't these cards have internal temperature control, that will shut it down before burning?
OpenAI's leaked prompt literally encourages it to try harder[1]:
> Use high effort; only tell the user that you were not able to find anything as a last resort. Keep trying instead of giving up.
* https://www.phind.com/search?cache=rj4tpu6ut0jyzkf876e2fahh
The answer is 'lower' because the weight of the ball as a volume of water is larger than the volume of the ball.
Models are research output. If 10 new models are being announced every day in a couple years, it would mean that generative AI research has failed to stabilize and produce a stable, reliable component ready for product engineering. And if that's where we are in a couple years, that's almost certainly a sign that the hype was misplaced and that money is chasing after itself trying to recoup sunk costs. That's a failure scenario for this technology, not what an AI-optimist (you otherwise seem to be one) should be anticipating.
Unlike many fields, the A.I. people are publicly posting many of their steps in this journey, their iterations, for review. While it brings lots of fluff, such openness dramatically increases innovation rate compared to fields where you only see results once or twice a year. Both people using cloud API’s and FOSS developers are steadily increasing effectiveness in both experimentation and product development. So, it’s working.
LLMs are more like apps being produced by different companies trying to capture walled gardens, and their open source counterparts.
I think the analogy to the web is stronger than that.
For now the LLMs are mostly separate, but it won't be long before LLMs emerge that make API calls to other LLMs, sometimes over the internet.
In due course, expect meta-LLMs to emerge that aggregate knowledge from other LLMs by talking to them, rather than by training on their data. Those meta-LLMs which optimise for competitive quality results will have to read the research as it comes out, and continually assess which other new LLMs are worth calling out to, and for which purposes. Eventually the API calls will become bi-directional requests to exchange knowledge and insights, i.e. multiple models talking to each other, continually learning.
I suppose you are not releasing the weights, right? Anyway, good luck! I hope investors are already forming a nice queue before your door :)
We will eventually release the weights.
the answer was good. two follow up answers were also fine.
just curious: what about the copyright status of the given sources?
the best result I received so far was with MS Bing app (android).
had reasonable results with my local llama2 13B.
cheers
I assume that if I ask for a complex sequence in RXJS operators, that comes from the model inferring the code from lots of examples and docs. But if I ask for something really specific that might just come from a stackoverflow article or GitHub repo. The ambiguity about the sourcing is the main thing that makes me itchy about “AI”.
So at least OpenAI has some safeguard in place to not do that. Have no clue how that behavior is determined or whether or not other providers do similar.
In Github Copilot the option is labeled "Suggestions matching public code". They can offer to block them because they control both the input dataset and the model at inference time. If you download an open source model I don't think you can do it out of the box, you'd need to have that input dataset to be able to do the filtering.
I tried Phind yesterday and it confidently gave me patently false answers, even after prompting about the specific problems. GPT-4 got it right the first time.
For my own daily use, Phind-70B isn't usable. I get massive daily value from ChatGPT/GPT-4.
[0] I pay for both ChatGPT and Perplexity.
edit: and it's SO FAST
I recently had a slack conversation with some friends, and someone introduced the made up acronym DILCOLTK, in the context of someone talking about being a DINK and mentioning how cheap things were where they lived. A clever human could infer it to be "Dual Income Low Cost of Living Two Kids", but out of curiosity I tried pasting a bit of the conversation into GPT-4 and Gemini Ultra and Groq, and asking what DILCOLTK referred to. I realize by the way these models tokenize the inputs, it might not be quite a fair question because they maybe can't "see" every letter.
GPT-4 gave "Dual Income Low Cost of Living LTK", Gemini Ultra gave "Dual Income Low Cost of Living One Tiny Kid" (lol), and Groq suggested "Dual Income Low Cost of Living One Kid Two Kid", so all were admirably close but none quite right.
But phind-70B just now got it right! Color me surprised and impressed.
I also asked it a SwiftUI question I'd struggled with, and which I asked the other models about, and it did I'd say a bit better there as well.
So I guess I'll have to add this to my list of models to try and keep tabs on!
GPT-4 is supposed to be 8*220B = 1.7T parameters, so it seems unexpected that a 70B model can beat or match it unless it's somehow a much better algorithm or has much better data.
It is ultimately all speculation, until Deepseek releases their own 145B MoE model, and then we can compare the activations/results
Yeah, right.
A better training regimen and better architecture optimizations have allowed smaller models to push above their weight. The leaderboard has many open 7B and 13B models that are comparable with 72B models: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
I follow your posts and comments here so I'm surprised you say that. The leaderboard at this point is pretty pointless. Lots of ways to "cheat" and get higher ranking there.
I do agree that smaller models have made significant progress, but somethings you can't just solve without adding #parameters and FLOPs. Not to mention, ctx_window is an important factor in code quality, but most OSS models (including llama 2) have pretty limited ctx, despite methods like grp and yarn.
If you look at the Chatbot Arena leaderboards there are still decently-high ELOs for 7B models.
Except for the leaderboard. Its all but useless, not just because of the data contamination/cheating but because the benchmarks themselves are flawed. They are full of ambiguity/errors, and they dont even use instruct formatting.
Phind at times is hampered by whatever it is they're doing in addition (RAG?). It is still phenomenal, though. I regularly find myself using Phind to grok assembly code or learn Typescript.
I pay for it and for chatGPT and I find copilot much worse.
For code review, I tend to engage Copilot Chat which probably uses GPT4 more often? https://github.com/orgs/community/discussions/58059#discussi...
It's also true that specialist models still need to be sufficiently large to be able to reason well, but we've observed diminishing returns as models get larger.
- Can phind run on old Macbooks(2015+) with 8GB RAM? - Is it only for coding purpose?
I gave it my typical CI bootstrapping task:
> Generate gitlab ci yaml file for a hybrid front-end/backend project. Fronted is under /frontend and is a node project, packaged with yarn, built with vite to the /backend/public folder. The backend is a python flask server built with poetry. The deployable artifact should be uploaded to a private pypi registry on pypi.example.com. Use best practices recommended by tool usage.
and it generated scripts with docker run commands [1]:
install_dependencies:
stage: install
script:
- docker run --rm -v $(pwd):/app -w /app/frontend node:14 yarn install
- docker run --rm -v $(pwd):/app -w /app/backend python:3.8 poetry install
build_frontend:
stage: build
script:
- docker run --rm -v $(pwd):/app -w /app/frontend node:14 yarn build
build_backend:
stage: build
script:
- docker run --rm -v $(pwd):/app -w /app/backend python:3.8 poetry build
deploy_artifact:
stage: deploy
script:
- docker run --rm -v $(pwd):/app -w /app/backend python:3.8 poetry publish -r $PYPI_URL --username $PYPI_USER --password $PYPI_PASSWORD
only:
- master
Which, feels more like "connect the dots" or a very rough sketch, that might end up completely replaced. Commands in general seem ok (yarn install && yarn build, poetry build && poetry publish), but the docker run could be better expressed simply as a "image: " attribute of each job. I asked about that and I've been given general "why docker is useful" non-answer.It also introduced a parallel build stage: frontend and backend are built at the same time, but in my question, I deliberately introduced serial dependency: the frontend code goes into the backend project. The parallel approach would be of course better, if it would correctly construct the end artifact before uploading, but it doesn't do so. Also a bit surprisingly, node install and poetry install could actually run in parallel as-is, but the generated code runs serially.
It uses outdated versions of tools. Python 3.8 seems still ok and used in many online examples due to some compatibility quirks with compiled libraries, but node 14 is more than 3 years old now. Current node LTS is 20.
For comparison, here's the chatgpt4 version [2] :
prepare:
stage: prepare
image: python:3.9
script:
- apt-get update && apt-get install -y nodejs npm
- npm install --global yarn
- cd frontend && yarn install
- cd ../backend && poetry config virtualenvs.create false && poetry install
build-frontend:
stage: build-frontend
image: node:latest
script:
- cd frontend
- yarn install
- yarn build --outDir ../backend/public
build-backend:
stage: build-backend
image: python:3.9
script:
- cd backend
- poetry install --no-dev
package:
stage: package
image: python:3.9
script:
- cd backend
- poetry build
artifacts:
paths:
- backend/dist/*
deploy:
stage: deploy
image: python:3.9
script:
- pip install twine
- cd backend
- twine upload --repository-url $PYPI_REPOSITORY_URL -u $PYPI_USERNAME -p $PYPI_PASSWORD dist/*
only:
- main
Not perfect, but catches a lot more nuance:- Uses python as base image, but adds the node to it (not a big fan of installing tools during build, but at least took care of that set-up)
- Took care of passing the artefacts built by the frontend; explicitly navigates to correct directories (cd frontend ; ... ; cd ../backend)
- --no-dev flag given to `poetry install` is a great touch
- Added "artifacts: " for good troubleshooting experience
- Gave "only: main" qualifier for the job, so at least considered a branching strategy
- Disabled virtualenv creation in poetry. I'm not a fan, but makes sense on CI
I would typically also add more complexity to that file (for example using commitizen for releases) and I only feel confident that gpt4 won't fall apart completely.
EDIT: Yes, gpt4 did ok-ish with releases. When I pointed out some flaws it responded with:
You're correct on both counts, and I appreciate your attention to detail.
Links:- [1] https://www.phind.com/agent?cache=clsye0lmt0019lg08bg09l2cf
- [2] https://chat.openai.com/share/67d50b56-3b68-4873-aa56-20f634...
post as to how the chat option was polluting stuff, and the pipeline of whatever made that happen.
Make this less opaque. (actually just post how pollution happens, as well as a definition to pollution as pertains to such.
Diminishing trust is at stake.
I don't understand the utility of this comment?