Step Away from Stack Overflow
elisabethirgens.github.io
elisabethirgens.github.io
I'll keep doing what I've always done:
Look things up on stack overflow (and other sites), compare different answers and check the discussions as well, figure out WHY people are doing what they do - and then do my own experiments based on that.
For me personally, this leads to faster learning results than experimenting completely on my own, without any guidance.
SO and its ilk give me an example that I can start to cross reference with the API level docs, to figure out how to pull it all together, and also where to look to find the "gotchas", tweaks, and options.
Example:
Compiler throws error 0x19822943 on what looks like valid code. Googling the error code along with the compiler name almost invariably returns stackoverflow results near the top, so I look at the post(s) with answers and see "Oh yeah this is a known issue, upgrade to version 1.29 and it goes away".
For all the issues I have with the site it is still pretty great for these sorts of problems that should be covered in the support documentation for whatever it is that failed (but often isn't -- or isn't in a way that google is likely to find it easily).
trying to find some kind of shortcut to just get my code working by copy pasta-ing random bits of spagetti and crossing my fingers. Stop. There’s a better way.
is about as honest as the Juice Loosener pitch from the Simpsons [0]. It's impossible to be able to just copy and paste something you find online and have it work, especially while ignoring all context and learning nothing.
This, a billion times.
I wonder if there is a way to voluntarily downrank pages in search engines so that answers that are obsolete do not show up first in results.
If I'm having issues with Ubuntu 22.04, the solution to a similar problem in Ubuntu 10.04 is very much likely to be wrong.
I took that away from the article as well, although I doubt that is the author's original intent.
One thing I like to do, if I have the extra time, is pull out a couple of different algorithm's books (Cormen always being one) and see if I can find a new way of viewing a problem in light of an "algorithmic approach."
Admittedly the research takes time and is always a lot of fun as I explore many other "rabbit holes" but it leads to other insights as well. The thing I will try to do is capture a few bullet points that I may have gleaned during this exercise. The approach is not conducive to goal, or time, sensitive deadlines, however.
Personally, I avoid SO because I've seen too many wrong answers. I prefer documentation, or blogs explaining approaches. But, it can be a resource, especially to new developers, to help you understand how you need to think and how to approach problems. As a replacement for coding it yourself, it's might be hot garbage, so you should understand it before committing to having produced that code and assign your name to it.
I do too, in theory. In reality though, 90 percent of the time it's essentially unnavigable / unsearchable / basically intractable to use.
What you want is simply something that tells you "how do I do X" where X is some real-world thing, like strip the end-of-line characters from a file. What you get from "documentation" as such is a huge sprawl of text explaining the guys of 90 different command line options. Good luck digging your way through that.
Imperfect though it is, SO was created to address precisely this huge, gaping disparity between what users want, and what most "documentation" actually delivers. Did I mention it's imperfect? And that you actually have to (shudder) think about the examples given, before randomly cut-and-pasting them into production?
The whole point is though, that it's better than nothing, and frequently is mostly correct (and it's not that hard to tell when the answer is wrong or requires a bit of fine-tuning for your use case). And at the end of the day, still saves me hundreds of hours compared to the nearly useless "documentation" that ships with most running languages and platforms these days.
Turns out, software matrix multiplication is not the best way of doing 3D graphics, but it's a great way of learning how the mathematics of it works. In the mean time of doing all this I had gotten linear algebra. LA was one of the first "hard" classes we took, and a lot of people barely passed. I got top grades in that class.
The moral of the story is that if you want to learn how something works, then this (what OP is doing), is correct. This is a great way to learn. There probably is a faster or better way of doing it, but in terms of learning-efficiency, this is great.
- Will Smith
This kind of exploration, banging your head against the wall, is exactly "beating on your craft".
No matter if the result is crappy by any metric, trying to wrap one's head around things and producing output - any output - over and over is oh so important.
And dare I say talent without practice is useless and a waste, because one's going to stay in the comfort zone, and usually get cocky about it, whereas relentless practice teaches humility.
s/art/literally anything/g in this comic:
Knowledge is about understanding how and why things work. That is what I was attempting to figure out about these linear algebra-operations. My classmates were attempting to retain the facts for the test, that typically doesn't work particularly well, and long term retention is poor. I've done that too. I could barely tell you what it is I'm supposed to have learned in those classes. But I could still probably construct 3D rotation matrices by hand without looking them up, all these years later.
Beyond that, what you describe as talent should be called built in ability, much like IQ or similar, because current usage of talent is too muddled with the well practiced.
Software matrix multiplication is a perfectly reasonable way of doing 3D graphics, when you don't have a GPU. They were just starting to become a consumer product at that time: https://fabiensanglard.net/3dfx_sst1/
Then a decade or so later I encountered matrices in computer graphics and suddenly I found that a lot of that stuff was actually useful and piece by piece it started to make sense to me.
The way that course was taught didn't work at all for me. I wasn't looking for motivation in the sense that I want to use it, but more about what was the motivation for developing any of it in the first place.
The line is somewhat blurry but there's probably a difference between playing a game and doing a job you don't really like because you're paid to do it.
https://rady.ucsd.edu/faculty/directory/gneezy/pub/docs/jep_...
ok, what is? all the 3d engines I've ever worked with do tons of software matrix multiplication. You generally need to know where things are for collision checking and other things so doing it all on the gpu is not an option
Sure. You can do some of it in software. That's fine.
Point is I didn't use the GPU at all, and did all the rendering in software. No matter how you rotate or translate it, that's not a particularly efficient way of doing it.
A series of calls to OpenGL is a heck of a lot easier than setting up 3x3 rotation and translation-matrices and implementing a library to do 3x1, 1x3, 3x3 matrix multiplications, scalar transformations, transpositions, calculating determinants, planar projections, perspective shifts, and what have you.
And as others have pointed out, a quick jaunt to Stack Overflow might have revealed to OP that leaning on the DB is what any seasoned engineer would do, which is probably a better engineering lesson than how to implement a reduce by hand.
AFAIK Stack Overflow's mission statement isn't "Be Rip-Offable for Devs Everywhere." Obviously I know about "copy-paste engineers" but that's a character trait, not the result of their tools. I've always treated SO more like a knowledgeable coworker to help when I get stuck.
Not too knowledgeable though, in my experience if you are the knowledgeable colleague in a domain, SO will not help you unless you're in one of the domains with a legendary SO contributor (e.g. C# because of Jon Skeet if you manage to get their attention), and unlike a knowledgeable colleague it won't assist you in looking for a clarification or solution either.
SO provides solutions to problems.
Experimenting, playing around and "throwing spaghetti at the wall" can lead to an understanding of the problem.
SO will tell me to use DISTINCT, and that is correct.
But why? Why is this the better? Is it really? Why is the home-brewed reduce slower? How could it be improved? Oh hey, but if I did this to the data upstream I could prevent non.distinct records from being in the DB in the first place, etc. etc.
Software Engineering is, among many other things, a craft. Just like carpenting or stone carving or smithing, it must be honed, and mastery can only be achieved if one plays, trys and experiments.
----
Of course, reading about the solutions to a problem can also help in its understanding. For me personally, if I do not understand a problem, I like to do both: Experiment, and read what the "canonical" way to solve it is.
But awkwardly dancing with the interpreter until something clicks in your head is highly valuable learning. In this case I kept waiting for the author to have a revelation or two - one about tuples, and one about proper use of dictionaries - but that never happened. I think it will in part 2, and then idiomatic python will fall out with the authors complete understanding of new concepts.
Then see if there's a better way to do it that looks smart.
Essentially, I'm agreeing with eska's comment, with the addition: "...but consider your data structures!"
groupby looks neat, but it requires sorted input, so it's very easy to convert an O(n) problem to an elegant but O(n log n) program.
EDIT: one could also do it with a one-liner in pandas, which means that as long as the input is less than ten million lines or so, the "import pandas as pd" statement is going to take longer than the actual program...
groupby_dict = collections.defaultdict(int)
for l in mylist:
groupby_dict[l['thing']] += l['count']
newlist = []
for thing, count in groupby_dict.items():
newlist.append(dict(thing=thing, count=count))
Knowing how do that does require the deep knowledge of data structures I'd expect every professional programmer to have, and doing a quick scan of the languages standard library (maybe an hours work) to see what it offers. After that you're set for solving not only this problem, but almost guaranteed to solve all problems you are likely to hit pretty optimally in Python. I have no idea where Stack Overflow fits into the picture.If you go a step further and do the Python standard tutorial found in the docs, you would discover how a finger weary experienced Python programmer might write those last three lines:
newlist = [dict(thing=thing, count=count) for thing, count in groupby_dict.items()]
Comprehensions are sweet Python syntactic sugar, but unlike mapping your existing encyclopaedic knowledge of data structures onto the languages standard library they aren't necessary for a casual user of the language. Indeed, some Python'istas will tell you resisting such delights make for clearer code.But trusting a quick Stack Overflow to tell you the optimal way to use a languages data structures - you must be kidding me.
groupby_dict = Counter()
for r in response:
groupby_dict.update({r['thing']: r['count']})
Counter is a subclass of dict, no further conversion needed.That's the key takeaway. Write it once to get it working and understand how to solve the problem, then write it better.
"Hence plan to throw one away; you will, anyhow." - Fred Brooks
There is very little record in general of the intermediate states people pass through on their trajectory from novice to expert. I suspect one could build an entirely new framework for education around such data.
That's exactly what motivated me to make a habit of going to SO first. But I don't just go there looking to copy and paste. I go there to learn and always review all the solutions given and the history and comments on them.
It was a struggle though because I was ornery and had made a habit of trying to figure everything out on my own. I generally know the logic required but struggled with the syntax to implement it.
Since I made that a habit I have learned so much from contributors there and made huge gains in productivity.
They don't really give a good or direct reason to "step away" from SO here though. It's really more about crafting code you can read and understand and that's good advice.
from collections import Counter
response = [
{
"thing": "A",
"count": 4,
},
{
"thing": "B",
"count": 2,
},
{
"thing": "A",
"count": 6,
},
]
counter = Counter()
for x in response:
counter[x["thing"]] += x["count"]
newlist = [{"thing": k, "count": v} for k, v in
counter.items()]
print(newlist)(This has to be said in a Yorkshire accent)
Particularly useful for frameworks like Svelte or React where I'm like... I have no idea how that's written.
If I was writing software professionally, I would think "this sounds like 'group by' in SQL, wonder what the equivalent is in Python", and then I would google for the the most common/idiomatic approach.
I would be interested if it's possible to do it in one-line using something from itertools.
import pandas as pd
[{'thing':x, 'count':y} for x,y in pd.DataFrame(mylist).groupby('thing' ['count'].sum().to_dict().items()] from itertools import groupby
keyfunc = lambda x: x['thing']
[{'thing': name, 'count': sum(t['count'] for t in things)}
for name,things in groupby(sorted(mylist, key=keyfunc), keyfunc)] import pandas as pd
pd.DataFrame(mylist).groupby('thing', as_index=False).sum().to_dict('records') var newlist = mylist.GroupBy(i => i.thing).Select(g => new {
thing = g.Key,
count = g.Sum(i => i.count),
}).ToList();you can learn a thing or two. people are genuinely here to teach. let's build a new SO because it's owned by china now.
So, pythonists, what’s the elegant answer?
d = collections.defaultdict(int)
for item in mylist:
d[item["thing"]] += item["count"]
newlist = [{"thing": key, "count": value} for key, value in d.items()]
(I'm assuming the order doesn't matter, and the specific output format is needed. E.g. if the next step then iterates over that list and unpacks it, obviously get rid of the last line and use the dict)There is the collections.Counter type, but I think there's not really a beneficial way of using that here.
Generally, the "there is one elegant solution" aspect of Python is widely overstated :D
I really thought if this was possible to do in a single dictionary comprehension, with some kind of d = {k: sum(v) for k, v in <something>}, but whatever goes in <something> ends up being super ugly.
[{"thing": k, "count": v} for k, v in [[c, c.update({x["thing"]: x["count"]})] for c in [Counter()] for x in response][0][0].items()]
SELECT a, b, SUM(c) FROM table GROUP BY a
doesn't even compile. You either project away `b` or group over both `a` and `b`. groups = itertools.groupby(sorted(mylist, key=lambda x: x['thing']), key=lambda x: x['thing'])
newlist = [{**group[0], 'count': sum(item['count'] for item in group)} for group in groups]
I'm not overly fond of it, having to sort the list for `groupby` is unpleasant and extracting values from dictonaries is verbose. If this was an array of tuples it could be made much more concise, but of course that doesn't allow storing extra information for each thing, which this solution does. def merge_counts(mylist):
thing_lookup = {} # map thing -> item
newlist = []
for item in mylist:
thing = item["thing"]
if thing in thing_lookup:
# seen before; update count.
thing_lookup[thing]["count"] += item["count"]
else:
# first time I've seen this thing
newlist.append(item)
thing_lookup[thing] = item
return newlist
It, like the original, mutates the original dictionary. I would prefer: # first time I've seen this thing
thing = thing.copy() # don't mutate the original
newlist.append(item)In javascript:
const mylist = [
{'thing': 'A', 'count': 4},
{'thing': 'B', 'count': 2},
{'thing': 'A', 'count': 6}];
const dict = mylist.reduce(({thing, count}, dict) => ({
[thing] : dict[thing] ? dict[thing]+count : count
}, {});
const newList = Object.keys(dict).map(thing => { thing, count: dict[thing] });
(Not tested or anything, so no guarantee that this works right off the bat, but something like this should work.) from functools import reduce
result = [
{'thing': k, 'count': v}
for k, v in reduce(
lambda result, item: {
**result,
item['thing']: item['count'] + result.get(item['thing'], 0)
},
mylist,
{}
).items()
] agg = {}
for item in mylist:
if item['thing'] in agg:
agg[item['thing']] += item['count']
else:
agg[item['thing']] = item['count']
result_list = []
for key, value in agg.items():
result_list.append(dict(thing=key, count=value))
In beginner land, reduce makes your eyes glaze over and dict.get(key, default) is understandable in principle but still confusing in practice.Edit: upon further reflection I think my application is at the tipping point where it needs another database table.
I can’t really remember it, so I’d probably go with defaultdict
E.g., every other language calls them "arrays", and uses `[ ]`, but Python calls (roughly) them "dicts" and uses `{ }`.
Every other language calls them "objects", and uses `{ }`, but Python calls (roughly) them "lists" and uses `[ ]`.
An Array is a contiguous section of memory with like items. Those like items _could be_ object pointers. The closest thing in Python to this is a tuple.
A List is a collection of things with an order, often a linked list or doubly linked list, but could actually be an array as well. Sometimes "array" and "list" are interchangeable.
A Dictionary (aka HashMap) is a key/value store. They are denoted in Python with a bounding {} (similar to like JSON objects). Unlike JS or JSON, the keys can be anything hashable, not just strings or ints.
I sort of get the impression that "every other language" means javascript, where we sort of abuse the syntax and treat objects as hashmaps (they aren't really) or arrays (also aren't really). Yet, Python maps quite well against JSON objects and lists.
Terminology definitely can be different from language to language, but I find Python to be pretty close to the C class of languages for that.
I have a nagging feeling there is an easier way to do this, but my quick and dirty solution was
def merge_list1(l):
other_dict = defaultdict(lambda: 0)
for t, c in ((i['thing'], i['count']) for i in l):
other_dict[t] += c
return ({'thing': k, 'count': other_dict[k]} for k in other_dict)
which is still readable, but probably far from optimal.Possibly not even then, it depends on how much you're doing and I feel like the topic at hand might be around that tipping point. We have some rather slow code that, profiling it, turned out to spend something like 60-70% of its time just converting between python types and native types when moving data in and out of the dataframe.