Events

Past Event

Data Science Day 2018

March 28, 2018
8:00 AM - 5:00 PM
Event time is displayed in your time zone.
Lerner Hall

Event Stats

  • 690+ attendees
  • 39 posters exhibited
  • 19 Interactive demos

Read our recap

Data Science Week Brought Together the Nation’s Leading Data Scientists

March 20, 2018

Columbia University hosted the first Data Science Week, which brought together researchers, industry experts, policymakers and other professionals working at the forefront of data science. Over the course of three daylong conferences (March 26-28), they explored how to harness the data science revolution in ways that enhance knowledge, benefit society and improve the quality of life.

“These three back-to-back conferences showed that Columbia is an international leader in data science,” said Jeannette M. Wing, Avanessians Director of the Data Science Institute at Columbia. “I’m delighted that prominent data science professionals from industry, academia, government as well as the nonprofit and foundation sectors came to campus to discuss how to advance the state-of-the art in data science and use data for good.”

Here are details on the three conferences that together constituted Data Science Week:

Data Science Leadership Summit

Data science leaders from colleges and universities across the U.S. gathered at Columbia for the Data Leadership Summit, the first meeting to bring together the academics who direct the nation’s data science centers and institutes. Some 65 heads of data science from more than 30 colleges and universities attended the March 26 Summit.

The Summit was organized by Wing, a former Corporate Vice President of Microsoft Research who is one of the nation’s foremost technology leaders. Since being named director of DSI in July, her goal has been to bring the directors of data science centers together.

“I’m excited to begin building an academic community in data science,” Wing said. “The overriding objective of the Summit was to explore what we can do together to advance the field of data science. It was a day where we shared best practices, discussed where we face similar challenges and opportunities, and talked about preparing the next-generation of data scientists as well as how to create a community around data science.”

As data science emerges as a discipline, data science initiatives at colleges across the U.S. and beyond have been sprouting up rapidly, added Wing. Whereas a few years ago there were a handful of data science institutes and centers, there are now more than 30 in top public and private universities.

Wing kicked off the Summit with a welcoming address titled “Data Science in Academia.” The morning’s speakers and discussions centered on data-driven research as well as the foundations and applications of data science. Participants discussed how to advance the state-of-the art in data science and how to use it to advance research in other fields.

The challenge for a university is how to support data science, an inherently interdisciplinary field, especially in its breadth of applications, when academia is often structured to excel within disciplinary boundaries, said Wing.

“Universities today answer this question differently – some create multi-disciplinary research institutes, some create new academic units while others add data science to existing programs or departments,” added Wing. “We are all in the early stages of figuring it out for our respective institutions. There’s no one right answer since each university has its own culture, traditions, and organizational structure. ”

The afternoon featured plenary talks, breakout sessions and roundtable discussions on topics such as ethics in data science, developing an academic cloud, best practices in undergraduate and graduate education and how to engage with industry, government and foundations to build a data science community. Kathy McKeown, the founding director of the Data Science Institute and a professor of Computer Science at Columbia Engineering, provided an overview on the National Academies Roundtable on Data Science.

Participants at the Summit worked in unison to answer these questions:

Research: How can we support inherently interdisciplinary research, if not in our foundations, then in our applications? How is faculty hiring (e.g., joint appointments) done? What kind of institutional support is needed? How are data science demands across campus being met (e.g., through applied data scientists, post-doc fellows)?

Education: What should every undergraduate know? What should every undergraduate data science major know? What should every master’s student know to be prepared for industry, not just the technology industry? What makes sense for courses, dissertation topics, and in terms of advising students? How do we advance the field of the data science and support its broad applicability at the PhD level

Ethics: How are people teaching students and encouraging research on ethics and privacy? What should we advocate as a data science community?

Engagement: What kinds of engagement do data science units have with industry, local government, foundations, other universities, and K-12? What is working and what does not work?

Compute Infrastructure: What kind of computing support is there for hosting data sets and providing data center-scale compute, including GPUs, etc., or do people use the cloud? How do we sustain open-source efforts that are germane to data science? How do we ensure secure access and sharing of data sets?

People: What kinds of personnel are data science units hiring to help bridge the demand for data science expertise across campus and in contending with the limited supply? Are universities creating a new kind of faculty or technical staff to hire these people? Are there new models that universities need to create to allow for a freer flow between academia and industry?

Sustaining a community: One intended outcome of this summit is to create an academic community around data science. How can attendees sustain such a community for the long-term?

The Summit was supported by the National Science Foundation, the Alfred P. Sloan Foundation and the Gordon and Betty Moore Foundation. Attending the Summit were leaders of the four NSF Big Data Hubs and the principal investigators for the NSF Transdisciplinary Research in Principles of Data Science (Tripod) awards. Columbia is a lead institution for both the Hubs and the Tripod research.

“The Data Science Institute at Columbia, with 300 affiliated researchers from every school and department at the university, is transforming all fields, sectors and professions through the application data science,” said Wing. “Columbia is a world leader in the field, and by bringing together data science leaders for this Summit, I know we will push the field a giant step forward.”

Annual Summit of the Northeast Big Data Innovation Hub

On March 27, The Northeast Big Data Innovation Hub, hosted at Columbia, held its annual summit that convened the data-science community of the Northeast United States.

The day featured updates on cross-sector initiatives, lightning talks from Big Data Spoke researchers and breakout sessions on data literacy, ethics, and health. The Hub, supported by the National Science Foundation, is a regional network of academic, industry and government partners who work in tandem to spur data-driven innovations and use data analytics to address society’s most pressing problems.

​“The Hub brings together data science leaders from across all sectors – academia, industry, government, and non-profit – to build a community greater than the sum of ​its parts,” said René Bastón, Executive Director of the Northeast Big Data Innovation Hub. “Our event showcased the breadth of perspectives among our data science ​stakeholders, and​ presented a great​ opportunity for ​them to collaborate on addressing ​large challenges with data-driven innovations.”

The Summit’s keynote speaker, Corinna Cortes, Head of Google Research, New York, discussed her team’s data-driven approach to fighting fake news. At Google, Cortes is working on a broad range of theoretical and applied large-scale machine learning problems. Additionally a panel of leaders from academia and industry discussed the challenges of rapidly advancing digital media, both in terms of maximizing its benefits and minimizing its potential drawbacks.

Following lunch, leaders of the Hub’s Big Data Spokes – multi-institutional, multi-sector collaborations that focus on topics of specific interest to the Hub community – highlighted the current work being done in their fields. Jane Greenberg, Professor at Drexel University and Director of the Metadata Research Center, talked about data sharing between sectors and the legal and privacy concerns that make sharing agreements difficult. Jaclyn Ocumpaugh, Associate Director, Penn Center for Learning Analytics, discussed the transformative effect that big data can have on education.And Chirag Lakhani, Research Fellow at Harvard Medical School’s Department of Bioinformatics, described a search engine (ExposomeDW) that finds environmental and phenotypic factors associated with disease and health.

There were breakout sessions in which stakeholders worked in groups to advance projects in data literacy, ethics, and health; a Data Literacy group worked to produce a draft framework of principles that define data-literacy concepts; and a Health group featured a talk from Vasant Honavar, Professor and Chair of Information Sciences and Technology at Penn State. There were also demonstration of the ExposomeDW given by Chirag Lakhani and Shreyas Bhave, an Undergraduate Research Assistant at Johns Hopkins. An Ethics panel featured lightning talks from a host of data science experts, which was followed by a group discussion and working session aimed at developing a proposal for a project on data ethics.

DATA SCIENCE DAY 

Breakthrough Research Featured at Data Science Day

In the final event of the week, Columbia held its third annual Data Science Day, a March 28 conference that showcased the research of university professors who use the most advanced techniques in data science to transform all fields, professions and sectors and solve some of society’s most vexing problems. The professors are affiliated with the Data Science Institute (DSI) at Columbia, a world leader in data science research, education and outreach. The day aimed to foster collaboration between data-driven innovators in academia, industry and government. Researchers also illustrated how they use data science to deepen their understanding of topics such as public health and medical research, climate change and financial risk.

“As a renowned university with top schools and departments, Columbia is the perfect laboratory in which to explore the transformation of all fields through the application of data science,” said Wing. “And for that transformation to spread widely, data scientists and domain experts must work with industry partners and policy makers to harness the data science revolution in ways that best serve society.  That was our purpose and hope in hosting Data Science Day.”

In quick summaries of their research, known as lightning talks, professors from across Columbia presented how they use data science to address problems like the onslaught of fake news and cyber attacks that threaten our privacy, our financial security and even our democracy.  The professors call upon the most advanced techniques in data science – machine learning, neural networks, topic modeling, Bayesian statistics and deep learning – to conduct their research. The focal point for this breakthrough research is the Data Science Institute, whose 300-affiliated faculty members work in all fields, departments and schools throughout Columbia University.

During a fireside chat at the conference, Wing discussed Data and Democracy with Columbia President Lee Bollinger. Along with being distinguished leaders, Wing and Bollinger are also prominent thinkers in their respective fields. Wing led Carnegie Mellon’s Computer Science Department and also oversaw the National Science Foundation’s computer and information science and engineering directorate. Bollinger is Columbia’s first Seth Low Professor, a member of the Columbia Law School faculty, and one of the nation’s foremost First Amendment scholars. In their chat, they discussed the implications of digital data on law, policy and democracy. They also discussed how platform companies such as Facebook, Twitter, and Google are heightening concerns about fake news and first ammendment rights. The two explored questions such as: Are there new threats to our democracy that we need to worry about? Or are we simply witnessing a transformation of what democracy means?

The day’s keynote address was given by Diane Greene, CEO of Google Cloud. She talked about the convergence of big data, artificial intelligence, and the cloud. Greene is one of the most prominent executives in enterprise technology. Before joining Google, she co-founded and sold three successful technology companies. She is on the board of MIT as a lifetime member of the MIT Corporation and remains on the board of Alphabet, the parent company of Google.

Here’s a list of the lightning talks given by Columbia researchers:

Health Discovery From and For Data Science

Panelists:

Nicholas P. Tatonetti, Herbert Irving Assistant Professor of Biomedical Informatics, and Director, Climate and Health Program; Using Electronic Health Records to Study the Heritability of Traits and Disease

Jeffrey Shaman, Associate Professor of Environmental Health Sciences; Developing Real-time Nowcasting and Forecasting of Seasonal Influenza

Jacqueline Gottlieb, Professor of Neuroscience: Knowing What is Important: How Humans Decide To What To Attend

Climate + Finance: Use of Environmental Data to Measure and Anticipate Financial Risk

Panelists:

Geoffrey Heal, Donald C. Waite III Professor of Social Enterprise in the Faculty of Business; Professor of International & Public Affairs: Evaluating the Risks of Sea-level Rise on Property Values and Coastal Infrastructure

Lisa Goddard, Director of International Research Institute for Climate and Society: Using Data to Help African Farmers and to Forecast Financing for Humanitarian Aid

Wolfram Schlenker, Professor, School of International and Public Affairs

Agricultural Yields and Prices in a Warming World

Machine Learning: The Good, The Bad and The Law

Panelists:

David Blei, Professor of Statistics & Computer Science:  Designing a Model to Show How Customers Choose Products

Junfeng Yang, Associate Professor of Computer Science: Effective Testing and Verification of Deep Learning Systems

Joshua Mitts, Associate Professor of Law: The Effect of Cybersecurity Breaches on Financial Markets

Professors and their students also demonstrated their research with poster presentations. The day ended with a reception for industry, government, faculty and students.

The Data Science Institute, founded in 2012, is an international leader in data science research, education and outreach. Its mission is to advance the state-of-the-art in data science; to transform all fields through the application of data science; and to ensure the responsible and ethical use of data to benefit society. DSI is also training the next generation of data scientists and developing innovative technology. With 300 affiliated faculty members working in all fields and schools across Columbia, the Institute fosters collaborations that advance the field of data science while addressing the urgent problems facing our society.

 

— Robert Florida

 

 

All speakers and their respected roles/titles are accurate to time of the event (2018)


2018 Keynote Speaker

photo of Diane Greene

Diane Greene

CEO, Google Cloud

Data, Insights, and Solution Journeys in the Cloud: This talk aimed to explore the convergence of big data, artificial intelligence, and the cloud.

Biography: After joining Google’s Board in 2012, Greene joined Google full-time in December 2015 as CEO of Google Cloud. Today, Google Cloud is one of the top cloud computing players, leading in data analytics and ML, agile open dev and deployment environments, security, and collaboration tools with G Suite. Prior to Google, Greene co-founded, ran, and sold three successful technology companies: VXtreme, a low bandwidth streaming video company, which was bought by MSFT; VMware, which EMC acquired and Greene took public for a $19.1 billion first-day closing valuation; and Bebop, an enterprise SaaS vendor acquired by Google. Prior to VXtreme Greene worked at two consulting firms as a Naval Architect, ran engineering for Windsurfer International, and worked at Sybase, Tandem, and SGI as a software engineer. Greene is on the board of MIT as a lifetime member of the MIT Corporation and remains on Alphabet’s (formerly Google’s) board. Greene served on the board of Intuit from 2006 through 2017.

Greene’s degrees include, an M.S. in computer science from the University of California, Berkeley, an M.S. in naval architecture from MIT, and a B.S. in mechanical engineering from the University of Vermont. Greene’s recent recognitions include being named to the Bloomberg 50, as well as receiving the Anita Borg Institute Technical Leadership Award, and a University of Vermont Honorary Doctor of Science, all in 2017. Greene is a lifelong sailor and was the 1976 Women’s National Dinghy Champion.



2018 Fireside Chat

With Lee C. Bollinger, President of Columbia University and Jeannette M. Wing, Avanessians Director of the Data Science Institute

Data and Democracy: A discussion of the implications of digital data on law, policy and democracy. Wing and Bollinger will touch on topics including how platform companies such as Facebook, Twitter, and Google are heightening concerns about fake news and first amendment rights, potential threats to our democracy and the transformation of what democracy means.

photo of Wing and Bollinger


2018 Lightning Talks

Session I: Health Discovery From and For Data Science

photo of Nicholas P. Tatonetti

Nicholas P. Tatonetti
Herbert Irving Assistant Professor of Biomedical Informatics, Vagelos College of Physicians and Surgeons, Columbia University

Talk Title: Disease Heritability using 7.4 Million Familial Relationships Inferred from EHRs

Abstract: Heritability is essential for understanding the biological causes of disease, but requires laborious patient recruitment and phenotype ascertainment. Electronic health records (EHR) passively capture a wide range of clinically relevant data and provide a novel resource for studying the heritability of traits that are not typically accessible. EHRs contain next-of-kin information collected via patient emergency contact forms, but until now, these data have gone unused in research. We mined emergency contact data at three academic medical centers and identified millions of familial relationships while maintaining patient privacy. Identified relationships were consistent with genetically-derived relatedness. We used EHR data to compute heritability estimates for 500 disease phenotypes. Overall, estimates were consistent with literature and between sites. Inconsistencies were indicative of limitations and opportunities unique to EHR research. These analyses provide a novel validation of the use of EHRs for genetics and disease research.

photo of Jeffrey Shaman

Jeffrey Shaman
Associate Professor of Environmental Health Sciences; Director, Climate and Health Program, Mailman School of Public Health, Columbia University

Talk Title: Nowcasting and Forecasting Seasonal Influenza

Abstract: In recent years, a variety of methods have been developed to estimate the current and future growth and spread of infectious disease outbreaks.  Here I describe some of the computational, mathematical and statistical approaches my research group has used to develop real-time nowcasting and forecasting of seasonal influenza.  Assessment of the operational accuracy of these nowcasts and forecasts will also be discussed, as well as ongoing efforts to validate and improve these systems.

photo of Jacqueline Gottlieb

Jacqueline Gottlieb
Professor of Neuroscience, Zuckerman Institute, Columbia University

Talk Title: Simplifying an Impossibly Complex World: Lessons from Biological Information Sampling Strategies

Abstract: Biological agents routinely make adaptive decisions in complex environments that they cannot fully comprehend. Faced with an overabundance of information, these agents have evolved a family of mechanisms for ignoring the vast majority of irrelevant inputs and very sparsely sampling relevant cues. Strikingly however, the question of how the brain generates sampling policies – how it determines what to attend to and what to ignore – has been until recently relatively neglected in neuroscience and psychology. I will review recent advances in understanding this question, with a focus on the emerging theme that sampling is not dictated solely by material gains but also by intrinsic factors including the uncertainty, effort and pleasantness of belief states that are expected to be engendered by the information. These cognitive forms of utility can be quantitatively characterized, and bring important insights into how intelligent agents cope with complex environments using potentially imperfect, yet efficient, question-answer strategies.  

photo of Suzanne R. Bakken

Moderator: Suzanne R. Bakken
Professor of Biomedical Informatics, Vagelos College of Physicians and Surgeons, Columbia University

 

 

 

 



Session II: Climate + Finance: Use of Environmental Data to Measure and Anticipate Financial Risk

photo of Wolfram Schlenker

Wolfram Schlenker
Professor of International & Public Affairs, Columbia SIPA

Talk Title: Agricultural Yields and Prices in a Warming World

Abstract: There is a strong nonlinear relationship between corn / soybean yields and temperature: yields increase in temperature up to roughly 30C (86F), when future temperature increases become harmful. The slope of the decline above the optimum is significantly steeper than the incline below it. Climate change has the potential to significantly decrease yields: the beneficial effect of shifting colder temperatures towards the moderate optimum is more than offset by the harmful effect of shifting moderate temperatures towards hotter temperatures. Changes in yields directly influence agricultural commodity prices, which are linked between periods through storage.  One third of reductions in agricultural production due to weather or biofuel mandates get offset by future supply increases, while the other two thirds come from reductions in demand.

photo of Lisa Goddard

Lisa Goddard
Director of International Research Institute for Climate and Society

Talk Title: Data & Finance in the Developing World

Abstract: Innovations in the creation and use of data in the developing world can help bring smallholder farmers out of poverty traps, and allow humanitarian organizations to move populations from crisis mode to risk management. This presentation overviews two examples in which enhanced observational datasets, seasonal  orecasts with reliable, quantified uncertainty, and cost-benefit analysis have been applied through collaboration with communities and decision makers to lessen climate shocks to the agricultural sector and make humanitarian aid go further. The first example is climate-based index insurance for smallholder farmers in Africa. The second example is forecast-based financing for the World Food Programme. The approaches can be complementary and also would be relevant to other sectors and geographies.

photo of Geoffrey Heal

Geoffrey Heal
Donald C. Waite III Professor of Social Enterprise in the Faculty of Business; Professor of International & Public Affairs, Columbia Business School

Talk Title: Rising Waters: The Economic Impact of Sea Level Rise

Abstract: By the end of this century sea level may have risen by anywhere between 65cm and 5m. The exact number depends on which models we use and what we assume about mitigation of greenhouse gas emissions. Even increases at the lower end of this range will have far-reaching economic consequences. I describe a project that is modeling these consequences. We bring together a database of residential property transactions in the US since 1980, flood risk maps, flood insurance premium data, LIDAR elevation data and data on attitudes towards climate change to assess whether exposure to the risk of flooding by sea level rise is affecting property prices. We model the behavior of property buyers in the face of flood risk and test this model on the residential property transaction database. We also evaluate the risks to coastal infrastructure associated with sea level rise. 

Moderator: Garud N. Iyengar
Tang Family Professor of Industrial Engineering and Operations Research, Columbia Engineering



Session III: Machine Learning: The Good, The Bad and The Law

photo of Junfeng Yang

Junfeng Yang
Associate Professor of Computer Science, Columbia Engineering

Talk Title: Effective Testing and Verification of Deep Learning Systems

Abstract: Machine Learning (ML) has made tremendous progress in recent years, achieving or surpassing human-level performance for a diverse set of tasks including image classification, speech recognition, and game playing such as Go.  These advances have led to widespread adoption of ML in security- and safety-critical systems such as self-driving cars, malware detection, and aircraft collision avoidance systems. Unfortunately, ML systems, despite their impressive capabilities, often demonstrate unexpected or incorrect behaviors on corner-case inputs, leading to disastrous consequences such as fatal collisions of self-driving cars.  In this talk, I’ll present some of our initial research towards testing and verifying the robustness of ML systems.

photo of Joshua Mitts

Joshua Mitts
Associate Professor of Law, Columbia Law School

Talk Title: Informed Trading and Cybersecurity Breaches

Abstract: Cybersecurity has become a significant concern in corporate and commercial settings, and for good reason: a threatened or realized cybersecurity breach can materially affect firm value for capital investors. This paper explores whether market arbitrageurs appear systematically to exploit advance knowledge of such vulnerabilities. We make use of a novel data set tracking cybersecurity breach announcements among public companies to study trading patterns in the derivatives market preceding the announcement of a breach. Using a matched sample of unaffected control firms, we find significant trading abnormalities for hacked targets, measured in terms of both open interest and volume. Our results are robust to several alternative matching techniques, as well as to both cross-sectional and longitudinal identification strategies. All told, our findings appear strongly consistent with the proposition that arbitrageurs can and do obtain early notice of impending breach disclosures, and that they are able to profit from such information.

photo of David Blei

David Blei
Professor of Statistics and Computer Science, Faculty of Arts and Sciences and Columbia Engineering

Talk Title: Shopper: Probabilistic Machine Learning for Consumer Choice

Abstract: I describe Shopper, a sequential probabilistic model of market baskets. Shopper uses interpretable components to model the forces that drive how a customer chooses products; it is designed to capture how items interact with other items. I describe an efficient inference algorithm to estimate these forces from large-scale data, and report a study of over five million transactions from a major chain grocery store. We are interested in answering counterfactual queries about changes in prices. We found that Shopper provides accurate predictions even under price interventions, and that it helps identify complementary and substitutable pairs of products.
This is joint work with Fran Ruiz (Columbia) and Susan Athey (Stanford).

photo of Adler Perotte

Moderator: Adler Perotte
Assistant Professor in the Department of Biomedical Informatics, Vagelos College of Physicians and Surgeons

 

 

 

 


 

Select Photos

DSI Industry Affiliates have access to Data Science Day posters after the event. If you are a current DSI Industry Affiliate, please please email us for a link to the videos.