2,704 karma · joined April 23, 2009
As I note at the end, the AI eliminated the knowledge bottleneck, but it still took 11 months of waiting for companies, agencies, and courts to reply.
Happy to answer questions about the routing or the EU261 technicalities. And to preempt the obvious one: a credit card chargeback only covers the original ticket cost. This was about getting the statutory €600/passenger penalty plus our overnight care expenses (and it was mainly the latter that prompted the whole story; Norse would probably have saved money if it had just arranged hotel rooms instead of handing us a $25 voucher).
I actually touch on this exact risk at the very end of the post. AI can automate drafting and knowledge retrieval, but that is only a fraction of the overall process.
The legal system is fundamentally slow, by design. It still took 11 months of waiting for agencies, airlines, and courts to move. That inherent slowness is creating friction to filter out low-effort "slop"; that latency partially accomplishes that.
Absolutely the easiest solution would have been to have a written exam on the cases and concepts that we discussed in class. It would take a few hours to create and grade the exam.
But at a university you should experiment and learn. What better class to experiment and learn than the “AI Product Management”. Students were actually intrigued by the idea themselves.
The key goal: we wanted to ensure that the projects that students submitted was actually their own work, not “outsourced” (in a general sense) to teammates or to an LLM.
Gemini 3 and NotebookLM with slide generation were released in the middle of the class, and we realized that it is feasible for a student to have a flaweless presentation in front of the class, without understanding deeply what they are presenting.
We could schedule oral exams during the finals week, which would be a major disruption for the students, or schedule exams during the break, violating university rules and ruining students vacation.
But as I said, we learned that AI-driven interviews are more structured and better than human-driven ones, because humans do get tired, and they do have biases based on who is the person they are interviewing. That’s why we decided to experiment with voice AI for running the oral exam.
For the use of LLM in classes: I understand the reasoning, but I found LLMs to be extremely educational for parsing through dense material (eg parsing an NTSB report for an Uber self-driving crash). Prohibiting students from using LLMs would be counterproductive.
But I still want students to use LLMs responsibly, hence the oral exam.
Such classes do not have the luxury of pen-and-paper exams, and asking people to go to testing centers is a huge overkill.
Take home exams for such settings (or any other form of written exam) are becoming very prone to cheating, just because the bar to cheating is very low. Oral exams like that make it a bit harder to cheat. Not impossible, but harder.
No, we do not want to eliminate the pen and paper exam. It works well. We use it.
The oral exam is yet another tool. Not a solution for everything.
In our case, we wanted to ensure that the students who worked on the team project: (a) contributed enough to understand the project, (b) actually understood their own project and did not rely solely on an LLM. (We do allow them to use LLMs, it would be stupid not to.)
The students who did badly in the oral exam were exactly the students who we expected to do badly in the exam, even though they aced their (team) project presentations.
Could we do it in person? Sure, we could schedule personalized interviews for all the 36 students. With two instructors, it would have taken us a couple of days to go through. Not a huge deal. At 100 students and one instructor, we would have a problem doing that.
But the key reason was the following: research has shown that human interviewers are actually worse when they get tired, and that AI is actually better for conducting more standardized and more fair interviews. That result was a major reason for us to trust a final exam on a voice agent.
Most banks will be risk averse and will not open an account to anyone applying online from abroad. Even for US persons applying online, they will ask quite a bit of documentation.
Some banks (but not all) will open an account for a non-US person, when the non-US person physically visits a US branch, with proper identification (typically a passport) and documentation on why they want the account. But even in such cases, it is up to the discretion of the bank employee to decide whether the risk of opening an account for a non-US person is worth the benefit. So, the same bank may give different replies to the same inquiry, depending on the branch asked.
As a concrete example, TD Ameritrade will open easily an account for a foreigner in the Chinatown branch in NYC, but will not open an account when the same customer visits a branch in midtown in NYC.
OpenTable has been a public company for almost 5 years now (see http://finance.yahoo.com/echarts?s=OPEN). Revenues, cost, growth, and all other metrics have been publicly examined and scrutinized for long time. The 46% premium paid by Priceline is based on how the new management estimates that they can leverage the assets of Opentable and hardly a "bubble-ish" premium.
If you believe that OpenTable is part of a bubble, then the whole US stock market is in a bubble, which may be true but again not directly connected to Uber and Whatsapp valuations.
The best example is "The five-dollar workday": http://en.wikipedia.org/wiki/Henry_Ford
Is greed limited?
Plus you cannot put a robots.txt at s3.amazonaws.com so if the url is accessed through the https://s3.amazonaws.com/.... url, the robots.txt will not work.
http://www.amazon.com/Race-Against-The-Machine-ebook/dp/B005...
There are many software applications today that are being seamlessly powered by a combination of human and machine intelligence. For example, Google Books is mainly digitized with OCR but then ReCAPTCHA is used to bring human intelligence for fixing the mistakes of the OCR process.
http://qmturk.appspot.com/ http://code.google.com/p/get-another-label/
You may find the code useful for what you are trying to do.
My point is that microtask work is not necessarily the optimal setting for tasks that are expected to last for longer periods of time. It is often beneficial to train and give people meaningful pieces of work instead of converting real work into micro-work and assume that workers are not intelligent enough to get things done properly.
I have to say though that there is a pattern: Once you solve and automate a process, people want more, and push you into doing a task that cannot be easily automated. Then you try to automate it, and the cycle continues...
For example, you want to create a caption for an image. You let a user create a caption. Then you take this caption and give it to another user, asking the user to improve it. Take the two versions and ask other workers, "which of the two versions is better?". Iterate until no improvement is possible.
Not a trivial setup, but gets around the binary accept/reject decisions problem and generates results of significantly superior quality.
I will just give here the answer that I posted in another thread http://hackerne.ws/item?id=2797371
the post was supposed to be "a story with the twist." Had I known that I was going to have hundreds of thousands of people reading the post, I would have followed the standard journalistic practice of writing a summary at the very first section. (See http://t.co/2kJEkJW for a copy of the post. Note: I asked the post to be taken down until I repost the original article but the journalist is really playing childish games.)
I kind of felt this a few hours after the post went out, so I added two clarification points early on in the article:
1. I am not giving up the fight, I will just fight differently, please see the conclusions (link),
2. This was not about NYU and people cheating in business schools; people cheat everywhere: the story gives an explanation why they remain undetected.
Oh well, people could not even read these two points.
But I want to write stories in my blog, not papers with an abstract, executive summary and table of contents.
Regarding the parable, I see the point and how what I wrote could be interpreted in the way that you describe. My intention was to refer to the city as the whole academic system. The "Redwich Village" is a reference to Greenwich Village, the neighborhood in NYC where NYU is located.
So, Redwich Village stands for my institution, which was attacked ferociously as being the harbor of cheaters over the last few days, while people fail to recognize that this is most probably a more generic problem. You may claim that I overgeneralize without proof, but I have sufficient evidence from other people that this is the case. Especially after all the emails that I received within the last few days where people described their own experiences with cheating that were very similar (most did so anonymously, but with specific names of top universities around the world).