HNHacker News
TopNewBestAskShowJobs

kat_rebelo

51 karma · joined August 13, 2020

submissionscomments
kat_rebelo··on The Inefficiency of Greed: How DeepSeek Exposed Silicon Valley's Tech Bros
there are reports that some openai employees initially learned about the release of the chatgpt interface via twitter. this move appears to have been orchestrated from on high by sam altman, who, despite his carefully curated public image, is not a scientist or researcher and holds no academic credentials at all, let alone any in the fields of computer science, machine learning, or linguistics. he is a prep-school educated kid who washed out of a comp sci degree at stanford that managed to pawn that off into being a VC who serves on the board of startups. in short, he is the exact kind of business guy this article is critiquing.

the release and viral adoption of chatgpt drove altman's personal profile and the valuation of the company he runs into the stratosphere, but, to many of the more sober/cynical minded of people who have been doing this kind of research for years (myself included), it appears to be at the cost of dropping a poorly understood (by the general public) technology with a high potential for abuse by multitude of different types of bad actors onto the general public with little or no plan for how to manage/mitigate the repercussions on the rest of society.

so yes, did openai play a "a significant role in accelerating the progress and acceptance of AI in our daily lives"... yes, but to many, that is not a good thing. we are only just beginning to scratch the surface of what this tech's impact will be on society as a whole. my guess is that most people with scientific / engineering backgrounds would have preferred a more incremental and controlled release process into broader adoption. instead, it seems like just another cynical move by another silicon valley pencil pusher relentlessly seeking to enrich themselves while accelerating the pace at which the billions of other people on this planet need to deal with downstream consequences of this action.

kat_rebelo··on The Theory and Technique of Electronic Music (2006)
this assumes the position that the pre-eminence of the even-tempered music based on european art music traditions and the associated staff notation. this is extremely limiting when considering the breadth of music that exists in the real world.

this is a theory of music, and while most pedagogy will reinforce the special position of this system, it is not THE theory of music. there are alternative systems of notation. there are harmonic systems that incorporate tones that do not exist in even tempered western scales. there are drumming traditions that are taught and passed down by idiomatic onomatopoeia.

this is especially apparent in electronic music where things like step sequencers obviate the need to know any western music notation to get an instrument to produce sound.

the western classical tradition is a pedagogically imposed straight jacket. its important to keep a more open mind about what music actually is.

kat_rebelo··on Netflix adds 9M new subscribers as it cracks down on password-sharing
this reeks of being a win for the pencil pushers at netflix who needed to turn around the slowly growing exodus from their platform on paper. i would guess that this will only be a one time boost as some people will have to get their own accounts. it seems like those people have already just done that.

it does nothing to address why there has been a slowly building exit from their platform... they have prioritized creating and promoting their own mostly-mediocre branded content instead of retaining licensing deals on movies and series people actually want to watch. all of this, while also continually raising rates.

kat_rebelo··on Can LLMs Reason and Plan?
LLM's only work on the data they have been trained on so all outputs are merely based on information that has already been written about by a human. Furthermore, LLM's do not truly "understand" even first order causal relationships, meaning whatever it plans will have no foresight to evaluate how a plan it generates will impact downstream components of a complex system.

LLM's live in "the world that has been written about", not the real world, and thus cannot formulate new ideas or hypothesis other than by accident. This, coupled with the lack of an ontological system for evaluating the validity of the statements it makes about a complex system, and, its lack of causal reasoning, means they cannot effectively plan.

I've worked on research related to causality that used LLM's (admittedly, pre ChatGPT and using much smaller models) and it was not uncommon to see extremely bogus causal relationships inferred such as "rising cost of living in NYC caused a flood in Argentina".

kat_rebelo··on Skip the API, ship your database
Even for more some more advanced use cases such as OLAP or ML related pipelines, in my experience, it takes a single senior developer a couple of days to design and implement a REST API, not weeks. That claim is overblown FUD.
kat_rebelo··on Fixing Penn Station
Its been a few years since I've had to go through Penn Station directly other than for LIRR. As I recall from my NE Corridor commuting days the lounge requires an Amtrak ticket. The NJ Transit waiting area had no lounge and no seats meaning people would begin to bunch up on the stairs. Admittedly, it might have changed in the intervening years.
kat_rebelo··on Marijuana addiction: those struggling often face skepticism
I am not sure about other drugs, but in the case of alcohol the rewiring is certainly not hyperbolic and not even limited to the brain. Humans have a special physiological relationship to alcohol and the human body will literally rewire both its neurological and digestive system to accommodate the increasing intake of long term alcoholics. It takes months to years for these changes to be undone, if they can be undone at all.
kat_rebelo··on G/O media will make more AI-generated stories despite critics
i have always augmented my more traditional music projects with avant-garde and experimental stuff. particularly focusing on found sounds, noise, and free improvisation. i record everything, but i treat it more as a journal of "free writing" than actual music output. the recordings are often formless, noisy, and chaotic and frequently contain experiments in polyrhythms or free time. over the years i've amassed an enormous amount of audio data that i treat more as an intellectual curiosity for myself than music i would show to other people.

however, with all of the debates around attribution and ownership for human creators in the age of AI art, combined with the apparently legally and ethically dubious means by which these megacorps obtain their training data, i have began to think about what are some avenues that the general public could try to protest the actions of these megacorps by discretely poisoning the well of their training data.

with the gigabytes and gigabytes of data i have generated of mostly incoherent audio, i have considered releasing this music for the first time by innocuously labeling it as a music audio dataset with the intention of trying to make it appear extremely attractive to megacorps scouring the internet for free data. my individual contribution probably couldn't amount to much, but if a concentrated mass of people did this in their respective fields, perhaps this could be a way of at least obstructing these corporations from freely capitalizing on the hard work of real artists.

kat_rebelo··on ChatGPT Explained: A normie's guide to how it works
i'm not a red state / far right / pro-trump in any sense of the word. however, i don't think it is very unreasonable to extrapolate bit and see the potential for societal harm.

over the past three years the entire world was impacted by a dire health crisis where misinformation played a large role in distorting public perception. this has direct impacts on public health (people not wearing masks, refusing vaccines) and has spillover effect into other parts of people's lives (political polarization around said issues).

what you see as a pretty cool toy could also easily be abused as a giant round the clock fake news generator. it doesn't matter if the text is true or even makes any sense... an alarming amount of people will take anything they read as fact without investigating the sources. this can be done as is with chatgpt right now. it has obscenity filters sure, but fake news is trying to pass as legitimate reporting, so it will probably be framed in a tone that escapes the obvious filters they have.

then consider the implications for robotexting, phishing, automated bots that pretend to be you to customer service chats, social media bots, messaging app scammers. all of these things are currently problems that can cause harm both personal and societal... and chatgpt will make it easier and cheaper to scale them up to new levels.

kat_rebelo··on Jailbreak Chat: A collection of ChatGPT jailbreaks
I could tell that this was generated by ChatGPT within two or three words. It's very funny that the link it selected for OpenAI's own ethical initiative leads to a 404.

Nevertheless, it failed to comprehend my point. I am not talking about ethical AI... I am talking about _auditable_ AI... an AI where a human can look at a decision made by the system and understand "why" it made that decision.

kat_rebelo··on Jailbreak Chat: A collection of ChatGPT jailbreaks
There is also the problem of causality. Humans are amazing at understanding those types of relationships.

I used to work on a team that was doing NLP research related to causality. Machine learning (deep learning LLM's, rules, and traditional) is a long ways away from really solving that problem.

kat_rebelo··on Jailbreak Chat: A collection of ChatGPT jailbreaks
The main reason is the mechanics of how it works. Human thought and consciousness is an emergent phenomena of electric and chemical activity in the brain. By emergent, I mean that the substrate that composes your consciousness cannot be explained only in terms of those electric and chemical interactions.

Humans don't make decisions by consulting their electo/chemical states... they manipulate symbols with logic, draw from past experiences, and can understand causality.

ChatGPT and in a broader sense any deep learning based approach, does not have any of that. It doesn't "know" anything. It doesn't understand causality. All it does is try to predict the most likely response to what you asked one character at a time.

kat_rebelo··on Jailbreak Chat: A collection of ChatGPT jailbreaks
No, chatgpt is based on a deep learning model where the core mechanics of the prediction involve millions (or billions) of tiny statistical calculations propagated through a series of n-dimensional tensor transformations.

The models are a black box, even the PhD research scientists who build them couldn't definitively tell you why they behave the way they do. Furthermore, they are all stochastic so its not even guaranteed that the same input will produce the same output, so how can you audit something like that.

This is a huge problem for many reasons. It's fine when its a stupid little chatbot, but what happens when something like this influences your doctor in making a prognosis? Or when a self driving car fails and kills someone. If OpenAI were interested in the _real_ social / moral / ethical implications of their work they would be working on something like that, but to my knowledge they are not.

kat_rebelo··on Manticore 6.0.0 – a faster alternative to Elasticsearch in C++
that is a slight misunderstanding of how open source licensing works.

the GPL bleed only happens if you distribute your application, meaning to sell or give away binary packages for customers to install. if your product is a hosted api that you do not distribute, you do not invoke that clause.

also, a lot of open source projects handle this by having things like the core engine licensed on a copy-left friendly license (GPL,AGPL). however, the language connectors and bindings are licensed under the slightly less restrictive apache license. unless you are offering a saas service of the product itself, it is more likely you are actually interacting with the connectors anyways. mongodb is a classic example of this model.

kat_rebelo··on Meta lays off 11,000 people
A/B testing is a completely different thing from hiring Behavioral Psychologists to design your platform to be as addictive as possible

https://www.goodreads.com/book/show/30962055-irresistible

kat_rebelo··on Too Many Songs, Not Enough Hits: Pop Music Is Struggling to Create New Stars
This article sounds a lot like major label dinosaurs and middle managers complaining that the business model that worked in the past doesn't work anymore. Can't say I have any tears to cry for them given that model really benefited labels and middle management instead of the musicians.

However, I don't really think that the music industry's woes are because of social media, viral breaks, or whatever they are attributing it to. If they didn't have their heads so far up their own asses it should be obvious to them why they are not breaking new artists. Every song in the top 100 approaches this new monogenre of R&B/trap/electronica that was clearly produced with GarageBand starter packs and autotune.

People who are passionate about music, buy records and merch, and go to concerts are probably more of the type of person who has niche interests and very deep and genuine love for a few artists / genres. They are not the kind of person who listens to the vapid musical fast food that major labels are shitting out.

kat_rebelo··on ZincSearch – lightweight alternative to Elasticsearch written in Go
being written in go does not imply anything about ease of deployment. kubernetes is written in go.

in fact, the type of project that sees something that works exceptionally well (like ES) and then decides to partially reimplement it (to the point of maintaining API compatibility) is a pretty big red flag that it is NOT going to be easier to use in the long run.

kat_rebelo··on We Are Changing the License for Akka
Infinispan and Red Hat being bought by IBM?
kat_rebelo··on Are we going back to the cable days?
for content you have purchsed (ie: prime) there is nothing stopping them from giving you a file in case they delist the content (Bandcamp is a good example of a service that offers exactly this).

this may be more of a political stance, but most of this content is produced by huge media corporations with deep pockets and histories of nefarious business practices. do i really care if disney gets paid for their content? no not really. if they amortize this loss by paying their employees less than that says more about them than it does about anyone pirating from them.

kat_rebelo··on Are we going back to the cable days?
i primarily torrented initially because i was broke. i was a super early netflix adopter (back in the DVD by mail days). for awhile netflix, hbo, and amazon prime were all i needed to watch 80% of the stuff i wanted to, and it felt fair to pay what i did.

with more content being paywalled behind the proliferation of superfluous streaming services, consistent rate hikes, and a huge drop in quality of content i've eliminated everything except prime (which i keep mainly for the two day shipping). i don't want to declare myself as a bellwether, but i got into netflix early and i got out early as well (cancelled my subscription Dec of 2021).

at the end of the day, torrenting is always the best value proposition. you can possess the files you download, they are not subject to being randomly taken down, and there are plenty of options for streaming them to multiple devices in your home. i download some really weird obscure stuff and i have never had an issue not being able to find subtitles. sometimes quality can vary but for most things you can get a decent 720 or 1080 rip.

kat_rebelo··on I want off Mr. Golang’s Wild Ride (2020)
Go is interesting because it has a cult like following of advocates who spout the same sound bites all the time, yet seem completely oblivious to the larger world out there.

Go seems to be in this weird space where it is not particularly suited for anything that its proponents say it is. As a systems programming language it gets a lot of flack for being garbage collected. As a web programming language, it is not ergonomic at all. The process of simply formatting a string is ridiculous compared to the string interpolation of Python or Scala. Furthermore, if I was a web developer, why would I want to deal with the cognitive overhead of Arrays vs. Slices or pointers?

The type system also takes a lot of heat, mainly for generics but that seems like something that is coming soon to the language. Nevertheless, it is sort of telling that Kubernetes decided to implement its own internal type system instead of leveraging the OO paradigm provided by the language.