Data Science Day 2020
Data Science Day 2020
Monday, Sept. 14, 2020
On Monday, Sep. 14, 2020, the Data Science Institute hosted Ethics & Privacy: Terms of Usage to bring together thought leaders driving the discussion around privacy and emerging technologies. As part of Data Science Day, the Institute’s flagship annual event, the event provided a forum for innovators in academia, industry, and government to connect.
Event Stats: Going virtual brought record-breaking attendance!
- 1,800+ Total RSVPs
- 780+ Live Data Science Day 2020 Viewers
- 1,500+ Overall Livestream Views
- 3,000+ Overall Web Engagements
Read our recap
Data Science Day 2020 Draws an International, Virtual Crowd
Sept. 14, 2020
More than 1,800 people from around the world registered for Data Science Day 2020, a virtual event during which Columbia professors and former Google CEO Eric Schmidt discussed ethics and privacy in data science.
In her welcoming address, Jeannette M. Wing, Avanessians Director of the Data Science Institute and professor of computer science, characterized data science as an emerging field that has profound societal consequences and “requires us to face [the issues of ethics and privacy] head on from the very beginning.” Data science relies on data, but data is about people, “about us, so how can we build models and systems while preserving the privacy of our data?” Wing postulated.
That was the central question that four Columbia professors addressed in short virtual presentations known as lightning talks. DSI member Tamar Mitts, assistant professor of international and public affairs at the School of International and Public Affairs, moderated the talks.
The NeuroRights Initiative: Human Rights Guidelines for Neurotechnology and AI in a Post-Covid World
Rafael Yuste, professor of biological sciences, discussed the NeuroRights Initiative, a mix of ethical codes and human rights directives he developed in an effort to protect people from potentially harmful neurotechnologies by ensuring the responsible development of brain-computer interfaces and similar neurotechnologies. To safeguard the development of all such devices, Yuste crafted five neurorights that can guide policymakers, technologists, and scientists who are regulating emerging neurotechnologies. Yuste has also written a technocratic oath, inspired by the Hippocratic oath, that he hopes all engineers, scientists and militarists working on neurotechnologies will sign and abide by.
“The goal of the NeuroRights Initiative is to preempt the creation of harmful neurotechnologies and AI algorithms by providing ethical frameworks for entrepreneurs, physicians, and researchers developing nanotechnology and AI,” Yuste said.
The Effect of Privacy Regulation on the Data Industry: Empirical Evidence From GDPR
Yeon-Koo Che, professor of economic theory in the Department of Economics, discussed his recent study on data privacy, which revealed an ironic twist: The European Union’s General Data Protection Regulation, intended to give people more control of their personal private data, had the effect of making it easier for advertisers to track certain people. In the two years since the General Data Protection Regulation was instituted, 10.7 percent of users opted out of sharing their data. But the trackability of remaining users increased by about 8 percent, according to Che’s study. He and his collaborators studied an anonymous third-party advertising company that recorded keyword searches and purchases for online travel agencies.
“We studied the impact of GDPR to highlight how government-mandated privacy protections interact with other privacy means,” Che said. “Do consumers benefit? That depends on how firms used improved prediction. Do firms suffer? They lose consumers from opt-out, but remaining consumers are of higher value to them.”
Data Science Ethics: A View From Public Health
How can data scientists, who work mostly with numbers, learn from public-health practitioners, who conduct human-subject research and are forced to confront the human element of data collection? In his lightning talk, Jeff Goldsmith, associate professor in biostatistics at Columbia’s Mailman School of Public Health, had a few recommendations for data scientists, such as to always consider consent, representation and the unintended consequences of their work.
He suggested that data scientists ask hard questions about data, such as who is included, is the measurement and sampling process valid, and what is the causal mechanism assuming or implying. Always involve stakeholders and plan for dialogue, Goldsmith said, and “understand who is likely to be impacted, and how, and engage at each stage—conceptualization, implementation, dissemination, which implies the need for transparency and openness.”
Differential Privacy: Basics and Latest Research
Differential privacy is a statistical approach used to protect the personal data of large groups. Here, individual records are intentionally left incomplete, while statistical properties are shared between data processing partners. Though it might seem like a sound way to safeguard data, people’s personal-purchase data can still be reconstructed with the release of accurate statistics from datasets, said Roxana Geambasu, an associate professor of computer Science who studies how to safeguard data used in machine learning.
She detailed how she and fellow researchers are developing ways to enhance differential privacy, which will have broad applications in helping safeguard personal data, she said. “We seek to address programming, testing and production challenges, provide tools for privacy budget management and incorporate different privacy into real infrastructure systems.”
Keynote Address: Eric Schmidt
After the four lightning talks, Eric Schmidt, executive chairman and co-founder of Schmidt Futures, discussed the intersection of human rights, data privacy, and AI technologies and what he has learned over the course of his career. Schmidt admitted that he has been wrong in some of his predictions about technology and economics, but said his errors were borne of an optimism about the use of technology for good.
The past few years, however, have shown the negative aspects of technology, and how to address that isn’t unambiguous or easy, he said. How, for example, do you regulate social media and fake news while safeguarding free speech? Or how do you regulate big tech companies without hobbling their creativity and growth? And how can the U.S. compete with China in the race for AI technologies, when the Chinese government handpicks AI companies to support and protect? These were some of the questions he discussed, first in his address and then in a fireside chat with Wing. He also took questions from the virtual audience by way of chat.
“I spent 40 plus years believing that technology was a strong force of good,” he said, “and I must say I get angry when the reality of the world collides with my relatively naive and simplistic view that technology should just make people better… But universities like Columbia are trying very hard to develop a broader and deeper understanding of where the technology really affects people.”
— Robert Florida
All speakers and their respected roles/titles are accurate to time of the event (2020)
2020 Keynote Speaker
Eric Schmidt
Former Google Chief Executive Officer and Executive Chairman and Co-Founder of Schmidt Futures
In conversation with Jeannette M. Wing, Avanessians Director of the Data Science Institute and Professor of Computer Science at Columbia University
2020 Lightning Talks
Ethics & Privacy: Terms of Usage
Yeon-Koo Che
Kelvin J. Lancaster Professor of Economic Theory, Department of Economics, Columbia University
Talk Title: The Effect of Privacy Regulation on the Data Industry: Empirical Evidence from GDPR
Abstract: Utilizing a novel dataset from an online travel intermediary, we study the effects of EU’s General Data Protection Regulation (GDPR). The opt-in requirement of GDPR resulted in 12.5% drop in the intermediary-observed consumers, but the remaining consumers are trackable for a longer period of time. These findings are consistent with privacy-conscious consumers substituting away from less efficient privacy protection (e.g, cookie deletion) to explicit opt-out—a process that would make opt-in consumers more predictable. Consistent with this hypothesis, the average value of the remaining consumers to advertisers has increased, offsetting some of the losses from consumer opt-outs. Joint research with Guy Aridor (PhD Candidate, Economics, Columbia University) and Tobias Salz (Assistant Professor of Economics, MIT)
Roxana Geambasu
Associate Professor, Department of Computer Science, Columbia Engineering
Talk Title: Security and Privacy Guarantees in Machine Learning with Differential Privacy
Abstract: Machine learning (ML) is driving many of our applications and life-changing decisions. Yet, it is often brittle and unstable, making decisions that are hard to understand or can be exploited. Tiny changes to an input can cause dramatic changes in predictions; this results in decisions that surprise, appear unfair, or enable attack vectors such as adversarial examples. Moreover, models trained on users’ data can encode not only general trends from large datasets but also very specific, personal information from these datasets; this threatens to expose users’ secrets through ML models or predictions. This talk positions differential privacy—a rigorous privacy theory—as a powerful, common foundation for building into ML much-needed guarantees of security, stability, privacy, and fairness alike. We draw upon our recent results using differential privacy to secure ML models against adversarial example attacks and to build a privacy-preserving Tensorflow-based platform that stops the leakage of training data through the models it pushes into production.
Jeff Goldsmith
Associate Professor, Department of Biostatistics, Columbia University Mailman School of Public Health
Talk Title: Building a More Ethical Data Science: Lessons From Public Health
Abstract: Goldsmith will discuss a public health perspective on ethical questions related to data science, which is shaped by a number of factors. In particular, research on human subjects imposes clear responsibilities, including respect for persons, beneficence, and justice, among others. More practically, public health researchers learn from data that can be messy and imperfect, and question whether an observation is a valid measure for a construct of interest, or if selection biases create a mismatch between the sample and target population. A shift towards complex data and analytic approaches makes these more critical than ever.
Rafael Yuste
Professor, Department of Biological Sciences, Faculty of Arts and Sciences, Columbia University
Talk Title: The NeuroRights Initiative: Human Rights Guidelines for Neurotechnology and AI in a Post-COVID World
Abstract: In my talk I will review the proposal that was made by the Morningside Group to introduce five new Human Rights into the Universal Declaration of Human Rights (1). These rights (“NeuroRights”) will protect mental privacy, personal identity, personal agency, equal access to cognitive augmentation and protection from algorithmic biases. Recently, protection of mental privacy has become particularly urgent, because of the fast development of brain-computer interfaces (BCIs) and the increased attacks to privacy due to COVID-related governmental measures. I will also review our proposal to follow a medical model, introducing a “Technocratic Oath” as a deontology in the Neurotech and data industry and using existing societal mechanisms similar to those already implemented in the medical industry to regulate future development of Neurotech and AI (2). Finally, I will discuss current advocacy efforts for NeuroRights in the US and different countries.
1. Yuste, R., Goering, S. and the Morningside Alliance Group (2017). Four ethical priorities for neurotechnologies and artificial intelligence. Nature 551, 159–163; 2017.
2. Goering, S. and Yuste, R. (2016). On the Necessity of Ethical Guidelines for Novel Neurotechnologies. Cell 167: 882-885.
Moderator: Tamar Mitts
Assistant Professor, School of International and Public Affairs, Columbia University
Posters & Demos
On Tuesday, March 31, 2020, the Data Science Institute hosted a virtual sneak preview of Data Science Day. This interactive gathering showcased 27 data science research posters and videos created by Columbia University faculty and students.
The Data Science Day 2020 virtual sneak preview brought together hundreds of remote attendees from around the world:
- 500+ RSVPs
- 700+ unique participants
- 300+ livestream viewers tuned in from more than 10 countries
The event has resulted in new partnership introductions and industry opportunities for its participating research teams. Further, these activities helped to bring more than 400 new people into the extended Data Science Institute community, who have signed up to learn more about admissions, student services, and upcoming events.
DSI Industry Affiliates have access to Data Science Day posters after the event. If you are a current DSI Industry Affiliate, please please email us for a link to the videos.