U.S. universities, rich in data, struggle to capture its value, study finds
newsroom.ucla.edu
newsroom.ucla.edu
Absolutely love it.
Let's distinguish a few things though. "Data science" seems like a pretty weird name. I mean, it's just "Science" right. Of course there's statistics, mathematics, signal processing, systems analysis, machine learning... all the good things that you and I are into.
But how does this get huddled uncomfortably beneath the umbrella "Data science"?
I think the answer is found by asking about the ends of data science, the old Cui Bono?
There's the raw entertainment value you mention. It's cool to have knowledge and visualise it. Sensors, transducers, processing is fun.
Then there's legibility. That is political and is about control.
What most scientists are doing with data is either hypothesis testing or combing for causal relations to then abductively feed back into hypothesis generation.
What most business people are trying to do is optimise, and adjust constraints and parameters. It's modelling for the most-part. It's ancient and goes back to linear analysis and regression from before the last century.
Security people are looking for stress signifiers, suspicious patterns with various triggers, selectors and tripwires.
Financial people want fortune telling. They want the models to extrapolate into beautiful hockey sticks.
Within any organisation we may need to do one, a few, many or none at all of the above. The problem then is that "Valuable data" is such a broad, open prospect it seduces gushing, credulous administrators into valuing the process, and the tools, but not the ends.
Instead we have social science, computer science and now data science.
Like you said, whether or not the visuals provide any value is a completely different line of discussion.
Have you looked at the Financial Modeling World Cup? https://www.youtube.com/channel/UCOlnCUAKLENyFC8wftR-oNw
Its tagline is: "Excel Esports. Yes, It's a thing"
For example: What is our retention rate? Meaning, what percentage of students who start our degree programs complete it. A fairly standard and important indicator of program health. Next, break this down by various cohorts: What is our retention rate among women? And so on. Heck, frequently we can't even answer questions about the current gender ratio within our program—and this is something that has been a focus of our diversity efforts recently.
I've had people say with a straight face that we _cannot_ calculate retention because we don't know when students leave our program. But of course someone knows this! And I've been able to produce rough estimates even given the limited data that I have access to. But a lot of educational data is fairly siloed, and frequently the people assigned to perform these tasks don't have much training and tend to give up quickly.
I suspect that many departments just don't have anyone assigned to do even basic educational data analysis on a regular basis, and with access to enough data to run interesting reports. My department is in the process of creating a faculty leadership role around academic data analytics, but my sense is that this will be a very unusual position. (And don't worry—it'll be filled by a faculty member, and not a new administrator.)
And don't even get me started about student evaluations of teaching. Yes, we give a survey at the end of every semester and ask students whether they liked a particular course and professor. No, those answers have very little to do with how much they actually learned. Yes, we could measure learning in other better ways—success in downstream courses, for example. No, people don't tend to do that.
There's a lot of room for improvement here, just working with the data we already have. No need for additional "telemetric signals".
Then you find out why retention is low.
Then you brainstorm ideas to increase retention.
Then you attempt to apply those ideas. It is at this point that the person responsible for applying those ideas says "been there done that".
Point is, most data dashboards are non actionable. The challenge is to create a good actionable dashboard (i.e. if values cross a certain threshold, then the user should take some action on it).
Once you create an excellent actionable dashboard, you realize it doesn't need a dashboard. It can be a notification.
So, while the data is important, the questions around it might just lead to the same work that was being done anyway.
The value in data driven approaches is high, but it takes a rare person to figure out why. Traditionally data has actually been a communication tool for things that are already known. That isn't at all how people expect it to be used, everyone seems to anticipate it is used to make decisions.
An organisation resisting data is bad news because it will struggle to talk about things that everyone knows to be true.
Huh? You don’t need a dashboard to make use of data. The best use of data in my mind is asking and answering questions.
Eg, maybe your program has a low number of graduating female engineers. Why? Maybe they’re dropping out along the way. Maybe female intake is low. With the data you can answer these questions.
The graduating rate of female CS students is low, but is it abnormally low compared to other schools? You investigate and - everyone has an equally low rate except one place where it’s 50/50. The data has led to a question - Why? What are they doing differently? And so on.
Maybe you find out that one specific class or set of classes is responsible for a lot of people leaving the program. So you zero in on that part of the curriculum and improve it. Sounds pretty actionable to me. We've actually done similar things on a smaller scale (week by week) to improve student success in my course.
But of course you don't know anything until you do the analysis.
The above are the type of questions that can result in different definitions of enrolled.
A similar problem in medicine: Every clinic system has a unique set of business processes, and a custom build of Epic. Granted the clinics are competing on which one can develop the most efficient processes, but does the patient benefit?
Of course not. What would benefit the students would be having a lot more standardization so that we compare ideas and approaches and determine what works. But the problem with standardized evaluation is that half of the programs suddenly discover that they aren't in the top half—as most of them had previously thought. This seems like more or less what happened to standardized testing in K–12 education.
But, in the context of a specific curriculum, having no idea what is happening and therefore no way to improve your bespoke curriculum is even worse than just deciding to do things your own way.
And it's worse than medicine, because at least they have some common metrics for what it means to be healthy. Whereas, faculty get to assign grades however they want! Imagine you ran a diet study where you both controlled the meals and got to reposition the numbers on the scale at will.
And anonymising data sources is a notoriously fraught task.
Just look at the treadmill of architectures designed to "solve" this problem: integrated databases to "data warehouses" to "linked data" to "data lakes" to "data fabric" to "data mesh."
Organizations that succeed do so because they have the resources to prioritize data management, and do so at significant team size and expense. But any organization that doesn't (or can't) view data curation as first class, top-line budget isn't in a position to capture even a small part of the theoretical value of their data.
Consider the "data mesh" architecture, the latest fad in this space. It basically assumes that not only do you have a dedicated data curation team, each data source also has it's own dedicated team of curators capable of productizing data sources.
That kind of thing is simply out of reach for most organizations.
A big part of the effort involved providing customers with big data solutions and all the implied benefits that would entail. They jumped aboard the hype train with gusto. Each quarter progress was reported and margins improved with a clear upward trend. From the trenches it was clear that even with the investment made, it was still not enough to reach any of the goals.
After a year of “modernization” the company was acquired, the buyer literally salivating over the coming profit margin increase and “synergies” in selling into related markets. After another year post acquisition it became clear that the benefits were simply out of reach, and the profits did not arrive. The acquiring company was left with the low margin business it had always been. Layoffs ensued as the buyer sought to salvage something from the transaction.
Turned out the entire modernization and big data effort had been an elaborate bait and switch committed by the board and CEO of the acquired company, masterfully executed.
Are you saying that rhetorically, or do you actually believe (perhaps with evidence) that the modernization and big data effort was not undertaken in earnest?
I feel like this part you omitted clarifies my opinion on the matter.
[1]: https://martinfowler.com/articles/data-mesh-principles.html
But the problem exists even on a concrete technical level. Look at how many different "big data" products are out there. The Apache project alone has probably 15 different tools that largely overlap.
The rush to "exploit" data reminds me of the dot com hype. It's one thing to use available data to make more informed decisions about things from course content to building occupancy. It's quite another to rush towards total surveillance because of a Fear of Missing Out of "exploitable data".
That's the primary value to be captured right there! Their privacy is highly valued, monetarily:
> authors contend that universities have been slower than organizations in other _economic_ sectors
Education in the US is just pure business.
https://www.projectcensored.org/ferpa-and-higher-ed-should-p...
It's a prophecy that demands to be fulfilled.
Starting with the premise "Data is useful", it proceeds to pick at all the ways we've failed to make it useful... and we're damned well going to make it useful if it kills us!
Maybe, just maybe (for those that dabble in the sceptical, explanatory game we call science) it might be that "bare data" has little use within certain contexts... say those that by definition are on the cutting edge of knowledge and reality, best steered by imaginative vision of leading experts.
Data driven market cybernetics may be very useful in some industries, such as ones with physical logistics, complex supply chains, rapidly shifting supply sources and demand sinks. But the primacy of "data uber alles" should not be one-size-fits-all.
I read a comment on here months ago about someone who enabled a hotel chain to process the data coming from the motion activated lights in every hotel. The chain was able to track workers with this data and cracked down on people taking too long of breaks.
And now a personal anecdote. The business I used to work at had an ID badge door you had to use to enter the smoking area. I was told by someone I trusted that when layoffs came, a member of management asked for a list of the people that used that door. The people on that list were prioritized for termination.
"Universities are literally awash in data. From administrative data offering information about students, faculty and staff, to research data on professors’ scholarly activities and even telemetric signals"
Why are professors' telemetric signals even being discussed? Are publicly listed office hours not enough? My dogs got chipped without their consent. Do professors deserve more dignity than my dogs?
Having data can undeniably be useful. The question is, how much ROI does the insights in the data provide, relative to the cost to build the organizational infrastructure capable of extracting these insights?
And there's a bootstrapping problem, because often it's hard to calculate the potential ROI without actually going through the work of building the infrastructure to process the data.
> how much ROI does the insights in the data provide, relative to the cost to build the organizational infrastructure
Having seen the workings of academia for a long time now I'd argue very little.
But I think the problem is more subtle. A Heisenbergian problem. When data collection involves people you have a triple problem of intrusion, distraction, and ossification. the act of trying to create a "data driven culture" in some contexts simply kills what's there. The cost of this, in addition to the bootstrapping costs of sensors and analytics are paralysing.
Academia, by its very nature, if properly functioning, demands to be dynamic. If it's good, it will not stand still long enough for you to look at it. And if you measure it "too hard" it will evaporate as all the good people who resent ossification leave.
There can be too much data. There are questions you cannot answer even by collecting all data, unless you have no time or food cost constraints but the value of your answer will decrease with the time distance from the moment the question was asked. In a million year you might know for sure who was communist in Atlanta during the 2000s, year by year, block by block, but you may care less by then.
There's no need to encourage them to start treating their mission as one of data mining in order to "capture" more value from the students.
I work as a BI dev at a large Big 10 school (50k students/35k faculty and staff). For all those students and staff, there are a grand total of 3 data engineers, who are responsible for all our central data warehouses, including DBAing our oracle and redshift DBs. Just to get a new (untransformed) table added to the data warehouse from a source takes 6+ months, and that's only if it's from a source with an existing integration.
On top of that, there are at least 50 people I know of whose job is basically to produce one or two reports manually in excel every week, and this isn't even considering people in finance or accounting. I'm talking about simple things like how many active research grants do we have, or how many students have enrolled in certain courses. These are reports (and entire fte positions) that could easily be automated with a single SQL query. Speaking of SQL, outside of the data engineering team, there are only 4 or 5 people who know any SQL out of the 100 I know of in reporting/BI positions. The "advanced" data teams are using MS access as an ETL tool to pull together data for tableau reports.
However, there are a lot of institutional issues that make fixing those problems difficult. For one, while we have a central IT dept, we also have about 10 individual college-level IT teams, which means that data isn't just in different databases, but on a whole separate network. For example, if I want to create a report on student faculty ratio, I need to connect to VPN 1 to export faculty data from redshift, then switch to VPN 2 to export student data from Oracle, then switch to VPN 3 so that I can upload both datasets to our depts SQL server. After all that I can finally write a SQL query to get a student faculty ratio. Oh and when we need to update that ratio in a month, I'll have to go thru the whole manual extract/load process again. Forget automating that, since the network teams have no incentive to allow any tunnelling or bridging from one network to another.
I could rant about this all day, but I think it's fair to say that there are still a ton of low-hanging fruit inefficiency-wise at universities. If we could get universities to value their data more highly, maybe that wod have the additional effect of solving some of these problems and even be a net money saver.
However, I note that your comments largely remark on an insufficient attention given to administration.
The trend in Universities over the past few decades has largely been to increase administration efforts, without that having a notable benefit on student outcomes.
But including security cameras in the list of tools seems… odd? What would they tell you? Classroom attendance I guess. But all that tells you is that the instructor is either so bad that nobody bothers showing up, or so good that their slides, recordings, etc are enough for the students to skip class sometimes. (And if you measure the number of students in the classroom, attendance will just become mandatory, this is dumb).
In general universities should be pretty inefficient I think. They are places where young people go to learn and try things out. Including student employees. If a university has no waste, that indicates that it hasn’t given enough possibly-unqualified undergrads enough resources to accidentally misuse.
> In general universities should be pretty inefficient I think. They are places where young people go to learn and try things out. Including student employees. If a university has no waste, that indicates that it hasn’t given enough possibly-unqualified undergrads enough resources to accidentally misuse.
The problem is all the extra money (yes, pretty much all of it) has gone to building up the administrations, not into resources that actually help students. And the students certainly haven't turned out smarter as a result.Universities have turned into a jobs program for people looking to work in education administration.
University data projects tend to succeed if and only if they're turned over to librarians, who are typically the only people on campus who have any clue how to do such a thing.
Universities are not an economic sector; they aren't profit-making enterprises. They are knowledge- and education-making enterprises.
Do they need more employees who aren't creating knowledge and educating? My understanding is that universities have been greatly expanding such non-core functions.
For example, everyone feels self-interest, but the reductionist claim it that it's all we are. A simple look at the evidence shows that it's manifestly untrue; people are much more than self-interested.
Similarly universities do and are much more. I know university students and they are learning a lot - I'm very impressed with the thought and creative effort put into the cirriculum.
The reductionists really hurt themselves and people who listen to them. They wall themselves off from all the good in the world, including education.
> I’m tired of discourse that leans towards “universities are places where we make you a complete human”.
I'm surprised that's said much. I see the reductionism as the mainstream; when I say otherwise, I'm 'politically incorrect'. On HN, certainly the reductionism is much more common. Isn't what you quote now the outsider view?
The same applies to corporations, landholders, monarchs and parliaments and governments.
New year is bringing out my fundamentals :-)
15750 administrators wanting to capture the value of data is just plain silly. (I’m going overboard here, but it’s about an order of magnitude of difference! I know of 50% differences in FTE/value in the market I work in, but x10…)
Also, it's worth noting that Stanford is in fact free for folks under the median household income in the U.S. (roughly speaking) which seems pretty 'affordable'. Of course the economics of one of these large-endowment, research intensive institutions is pretty much unrelated to their teaching function. But that just highlights the weakness of using gross, whole-institution numbers of people/dollars for any sort of comparison. Big universities serve lots of ends (not just teaching undergrads) and so the teasing apart the economic picture (and whether it's efficient at meeting it's many goals) is complicated.
This extends through a huge range of administrative functions for everything from calling snow days to collecting taxes to pay for the school etc.
Yes, absolutely, universities are awash in data they do little to anything with outside of hyperfocused laboratories. I wonder if having central data management groups would be an enabler or a bottleneck, however.
I think this article rings true for most universities, but for some, like mine, we’ve had unified systems for the last 20 years. While that’s been helpful in a lot of ways, the primary thing we’re looking for - helping students graduate and get better jobs than they have today - is surprisingly challenging. Adult learners often fail to finish school for a variety of reasons, most of which aren’t academic and often can’t easily be forecast using all available data.
Unifying data sources or mining them further doesn’t solve student health issues, childcare, jugging work responsibilities with school, etc.
What’s more, universities are slow-moving, heavily bureaucratic, and highly regulated. So even when we have good ideas to relieve burden those often come head to head with accreditation rules, financial aid regulations.
In my opinion, for my students at least, data is the least of our problems.
They have 0 visibility into any of the operations and therefore can’t derive any insights or install controls.
They have no idea if:
Some professors more effective teachers than others.
Which days students buy more food.
Whether taking a given prerequisite leads to better mastery of advanced material.
If additional spend is correlated with better student outcomes.
For people in tech businesses, it’s such a shocking thing to read it’s hard to fathom how any of the leadership at universities find this situation acceptable.
Teaching outcomes are evident from exam results, grants won, papers published etc and are routinely used to determine promotions, staffing and more. Food/footfall etc are monitored and predicted by the private catering companies on campus. Budget reviews are comprehensive and require yearly and greater analysis of university performance on a per department basis. All that doesn't even include competitive processes like ranking tables for universities, applications for grants, participation in international research groups etc etc.
What are you even talking about? Literally every course at most universities end with an evaluation. Obviously, there are other problematic ways to measure effectiveness (publications). How much that data gets used is another question and varies widely by institution. “People in tech businesses” love to exaggerate the quality of the data they have and its usefulness, hence the Subprime Attention Crisis that is rapidly eroding their cash on hand as the Fed’s cheap money and QE dries up.
I’m in a tech business and it doesn’t sound shocking at all. The only things people really have data for is how many people are giving them how much money.
1. Systems in higher Ed often do not talk and departments tend to be siloed and have drastically different goals and objectives.
2. There are a lot of privacy concerns as most of the data is students data.
In addition to this, I question what the end goal of the data analysis is? What is the value for a not-for-profit? The goal is going to be based on the department as a college doesn’t have the same profit motive. The individual departments typically don't have the resources to do deep data analysis.
I don't see a single example of "success" in the article. (They said "We unexpectedly found a pervasive void of infrastructure thinking and a relatively limited set of data-informed planning successes" but I must have missed the successful example.)
So the entire article describes what they think should be done, without any successful examples to show that anything useful can be done in the first place.
I had the same thought - intrusive analytics and a focus on exploiting and monetizing client data seem to have damaged the web, made gaming less fun, and made everything user-hostile and creepy. Do we really want universities to go down that path?
Or: This lecture brought to you by Pepsi.
I personally think this is wrong. Big tech companies should build hardware, while government agencies should use that hardware to curate our data. Like how it was near the beginning of the internet. Companies should be as far away from our data as possible.
Stop running dragnet data collection on unwilling participants and stop using it to try to “optimize” your interactions with them.
This data shouldn’t exist, it’s value shouldn’t be exploited.
Perhaps we could actually do something about this if we focused strongly on falsification, but right now there's just going to be a whole lot of Wittgenstein's ruler going on; what people want to look at or see will come first with such an abundance, and there's going to need to be some kind of real filter to make it useful.
The funniest part was the example of security cameras. What big breakthroughs are universities hoping to achieve with this data?
This 4 parter I wrote for Techrights [1] and the Times article that seeded it [2] go into detail.
[1] http://techrights.org/2022/12/28/andy-farnell-on-british-uni...
[2] https://www.timeshighereducation.com/campus/eliminating-harm...
I am absolutely sure that donor data is organized, systematized, useful and generates positive ROI.
It’s surprising that this study would not mention donor data (unless they had a predetermined “struggle with data” point of view in mind)