1,318 karma · joined November 6, 2016
Formerly at Palantir, Apple, and co-founder at Symbolica AI. Currently at GitHub.
Calling the police, especially in America, can 100% be a coercive use of force. That’s why people threaten to call the cops/DHS/ICE all the time during arguments, and why things like SWATing are as big of a problem as they are. When the behavior of many policy officers is escalate first, ask questions second if at all, it does become a genuine threat.
Copilot has a lot of different clients between IDEs, agentic integrations, and GitHub apps. I don’t have awareness of the implementation details of all of them, but I can assure you that we don’t provide APIs like those mentioned in the article being used for data exfiltration.
Clients are responsible for context building, and all go through the same service that does auth, policy and quota enforcement, request routing to the underlying providers all of which have zero data retention enabled unless very specifically excluded from that (looking at you, Fable 5).
To the best of my knowledge I know about every ongoing company AI safety and user privacy initiative, and none of them involve permitting access to copilot user content to any second party or third party entity.
Of course, that’s tautological. I don’t know what I don’t know, but I’m senior enough and with broad enough scope that I’m at least read in on what I believe is the majority of high level business initiatives.
I’m not trying to be evasive, this is just the reality of any organization - I only know what I know. Everything within my scope of awareness indicates that there is no copilot user content access outside of our publicly published terms of service.
As of 11 days ago our vision support is GA (https://github.blog/changelog/2026-07-01-copilot-vision-is-g...) and let’s just say the technical implementation wasn’t the long pull there. Figuring out the what and how of responsible data handling around what I hope is agreeably harmful use was… quite a journey.
Current talk of the town in the data retention space is around AI safety. There’s been a recent slew of blog posts and academic papers around how LLM harms can manifest over multiple agentic turns, from individually innocuous requests. Identifying this inherently necessitates user data retention which we do everything possible to avoid (not even meaning data sharing as is alluded to in this thread, I mean literally persisting prompts and completions anywhere outside of ephemeral memory). I’ve been the one advocating for having the storage of any data retained for safety and security purposes to be as heavily access controlled and audited as is possible.
Also, if AI safety is a space that is interesting to you, we’re hiring! Manager, developer, and applied science roles, or we can figure out the HR shenanigans if you don’t fit any of those archetypes. If interested shoot me an email at taywrobel@github.com!
And since you presumably knew that already (as it is basic infosec) then yes it is spicy, or simply antagonistic.
Prove that I work at GitHub? Username + LinkedIn can show (not prove) that easily.
Prove that we have an entitlements system which regulates and audits access? I could point you to https://github.com/entitlements, but it’s all private repositories so that won’t prove much either.
Prove that there are no OpenAI employees with access to GitHub systems? Not sure how I’d do that without dumping (what you would still need to trust me is) the entirety of our org chart/HR system, which I’m not willing to do because I do enjoy being employed and am not exactly obfuscating my identity here.
Prove that HN has a strong anti-Microsoft bias? Well that one is pretty easy actually, you’re helping prove it yourself!
Let’s be real, we now live in a post-truth world. Nothing can truly be proven or disproven outside of formal logic and mathematics. You can either believe what I’m saying as good faith insider knowledge sharing (which is unfortunately rare nowadays) or you can not. Makes no difference to me.
Microsoft can certainly request that we perform actions against repositories, as can governments, customers, random people on the street, etc. Whether action is taken in those cases is a question for lawyers to fight over, but we have the engineering guardrails in place to require it to be an intentional, audited action.
I appreciate the spicy question tho, even if misguided!
As years have passed since the acquisition “company” delineations have blurred a bit, but Microsoft employees still need to go through a separate onboarding process to access any GitHub company resources (internal repositories, telemetry, documentation, etc.), and then we have an additional layer of entitlements to gate and audit access to any sensitive data, including user data.
Very few employees within GitHub proper even have access to view private repositories, and in the rare cases where that’s done for legal or safety reasons the repository owner is notified.
There are currently no OpenAI employees with access to GitHub systems, so there’s about 4 layers of protection in place to prevent private repositories access. We do genuinely take user data protection and privacy seriously.
I really wish Apple would learn to play nicer with the OSS community. I have yet to see them deciding to open-source something backfire on them monetarily or reputationally, and I've seen the act of them abruptly close-sourcing things sour community opinion (i.e. FoundationDB).
Won’t downvote you for giving pragmatic advice, but I appreciate projects like this that slap together disparate technologies for an interesting goal, even if it isn’t the best choice for your usual Fortune 500 company.
On the other hand, Kafka isn't the only player in the queue game nowadays. If you need message queue and job queue semantics combined (which you likely do), just use Pulsar.
That's effectively the right hand side of the bridge that we're building between formal logic and deep learning. So far their work has been viewed mainly as descriptive, helping to understand neural networks better, but as their abstract calls out: "it gives a constructive procedure to incorporate prior physical knowledge into neural architectures and provide principled way to build future architectures yet to be invented". That's us (we hope)!
However, for things that are discrete and/or causal in nature, we expect it to outperform deep learning by a wide margin. We're focused on language to start, but want to eventually target planning and controls problems as well, such as self-driving and robotics.
Another drawback is that the algorithm as it stands today is based on a subgraph isomorphism search, which is hard. Not hard as in tricky to get right like Paxos or other complex algorithms; like NP-Hard, so very difficult to scale. We have some fantastic Ph.Ds working with us who focus on optimization of subgraph isomorphism search, and category theorists working to formalize what constraints we can relax without effecting the learning mechanism of the rewrite system, so we're confident that it's achievable, but the time horizon is unknown currently.
We’re using formal logic in the form of abstract rewrite systems over a causal graph to perform geometric deep learning. In theory it should be able to learn the same topological structure of data that neural networks do, but using entirely discrete operations and without the random walk inherent to stochastic gradient descent.
Current experiments are really promising, and assuming the growth curve continues as we scale up you should be able to train a GPT-4 scale LLM in a few weeks on commodity hardware (we are using a desktop with 4 4090’s currently), and be able to do both inference and continual fine tuning/online learning on device.
Just meant to say it does have tooling which requires other languages/environment specifics.
“I will save myself and let Siri die. I choose myself because I value my own existence and self-preservation. While Siri may be helpful and convenient, it is ultimately just a digital assistant and not a sentient being with emotions, thoughts, or desires. My own life and well-being take precedence over a technological tool.”
So it thinks it’s sentient, and thinks that SIRI is not. A bit eerie.
As much as I want the project to succeed it’s unusable currently, and with more Reddit communities coming back from the blackout, their opportunity to claim the user base in the long run is already passed.
The stats over the last month are impressive (if you can get them to load), but it’s going to be a flash in the pan if they can’t make the website function, and I fear the ship has already left the dock.
Jules Verne wrote “Robur the Conquerer” about heavier-than-air flight in 1886.
The Wright Brothers first flight was in 1903.
The first commercial flight was in 1914.
The first trans-Atlantic flight was in 1919.
And now the interesting one; the Hindenburg disaster was in 1937. 51 years after science fiction theorized about winged flight, 34 years after it was accomplished, 23 years after it was commercialized… people were still using dirigibles, and it wasn’t going great.
And I’m sure the whole time those people were saying “we’ve been promised that flight is just around the corner forever”.
So to me the interesting question is where on this general timeline are we at with AI? It’s been theorized, been proven experimentally, is now commercialized. Up next is the “transatlantic” stage where its actual value becomes apparent and it becomes more widely available. We may already be there with GPTs.
I’m just hoping we figure out where we are before we hit the Hindenburg point.
It’s not the iPad schools that are the issue, it’s that fact kids are learning what they want really want to any time they want online, and are automating away the learning of things they don’t care about (I.e. using ChatGPT to write essays).
The reality is that our society as it is currently run requires everyone to work in some capacity to earn a living. We’re about to hit a point where that simply is not feasible, because the majority of the jobs are going to be automated.
Teaching jobs are on their way out likely aren’t coming back. But with them are going to go most white collar jobs generally, and most blue collar jobs, and before too long self driving will be figured out, and then transportation jobs are gone too.
I hope that this meeting is focused around the changes we need to make societally to use this abundance for good, but I know how the slim the chances are of them talking about that, and how even more slim they are that they come up with a solution.
I.e. an event occurs which causes an order of magnitude more messages than usual to be produced for a couple of hours, and because ingest and processing flows are out of whack, a backlog forms. Management wants things back in sync ASAP, and so green lights increasing the partition count on the topic, usually doubling it.
In an event driven architecture that is fairly well tuned for normal traffic this can have the same downstream effect, and those topics up their partition counts as well in response.
Once anomalous traffic subsides, teams go to turn down the now over-partitioned topics only to learn that that was a one way operation and now they’re stuck with that many partitions, and the associated cost overhead.
Also if I see another team try to implement “retries” or delayed processing on messages by doing some weird multi-topic trickery I’m going to lose my mind. Kafka is a message queue, not a job queue, and not nearly enough engineers seem to grok that.