Data Science Day 2023
Data Science Day 2023
Wednesday, April 19, 2023 (8 a.m. – 5 p.m.)
Data Science Day provides a forum for innovators in academia, industry, and government to connect. The April 19, 2023 event featured a keynote presentation from Manuela Veloso, Head of J.P. Morgan Chase AI Research; and Herbert A. Simon University Professor Emerita at Carnegie Mellon University; three sessions of Columbia-led lightning talks; interactive posters and technology demonstrations; and remarks from Lee C. Bollinger, President of Columbia University. Clifford Stein, Interim Director of The Data Science Institute; Wai T. Chang Professor of Industrial Engineering and Operations Research and Professor of Computer Science, was the master of ceremonies. 2023 marked the in-person return to Alfred Lerner Hall on the Morningside campus.
Event Stats
- 500+ attendees
- 85 research projects exhibited, including 79 posters and 6 demos
2023 Keynote Speaker
Manuela Veloso, Head of J.P. Morgan Chase AI Research; and Herbert A. Simon University Professor Emerita at Carnegie Mellon University
Manuela Veloso is head of J.P. Morgan Chase AI Research and Herbert A. Simon University Professor Emerita at Carnegie Mellon University, where she was previously Faculty in the Computer Science Department and Head of the Machine Learning Department. Her recent interests are in Artificial Intelligence (AI), Symbiotic Human-Robot Autonomy, Continuous Learning Systems, and AI in Finance. She is past President of the Association for the Advancement of Artificial Intelligence (AAAI), and the co-founder and a past President of the RoboCup Federation. In her career she has received numerous awards and honors, including: National Science Foundation CAREER Award, Allen Newell Medal for Excellence in Research, Radcliffe Fellow, Einstein Chair Professor of the Chinese Academy of Sciences, and the ACM/SIGART Autonomous Agents Research Award. Veloso is a Fellow of AAAI, AAAS, ACM, and IEEE. She was elected in 2022 to the National Academy of Engineering for her “contributions to artificial intelligence and its applications in robotics and the financial service industry.”
Talk Title: Symbiotic Human-AI Interaction: Experience-Based Insights from the Finance Domain
Abstract: In this talk, I share insights on the interaction of humans and AI to jointly solve end-to-end complex problems in the financial domain in particular. I will focus on data discovery, data standardization, synthetic data generation through simulations, data reconciliation, and explainability. I will present the challenges and opportunities for a symbiotic interaction between humans with their principles, knowledge, and experience, and AI with its ability to learn from data and from feedback. The talk will include a discussion on the use of LLMs in multiple tasks.
Moderated By: Jeannette M. Wing, executive vice president for Research and Professor of Computer Science, Columbia University
Presidential Remarks
Lee C. Bollinger, President, Columbia University, joins the event to give remarks on the impact of data science and the Data Science Institute. Bollinger will be joined on stage by Clifford Stein (current DSI Interim Director); Jeannette M. Wing (former DSI Avanessians Director) and Kathleen R. McKeown (DSI’s founding Director), for a special ceremony to recognize his presidency and his contributions in the establishment of the Data Science Institute.
Photo Credit: Eileen Barroso
2023 Lightning Talks
The Human in AI Systems
Kelton Minor
Postdoctoral Research Scientist, Data Science Institute, Columbia University
Talk Title: Global Monitoring of Emotional Responses to Climate Extremes: Evidence from Eight Billion Social Media Posts
Abstract: Climate change is intensifying regional heat and precipitation extremes, posing complex risks to human well-being on a planetary scale. Can pairing digital data streams with NLP provide a tool to track the hidden human impacts of climate stressors on daily life? In this talk, I’ll share key insights from a global-scale natural experiment that linked the lexical content of ~8 billion geolocated tweets across 190 countries and 13 languages with daily data on local climate extremes and weather conditions. Constructing historical sentiment atlases for nearly every county in the world, I’ll assess whether local exposure to randomly-timed climate hazards alters positive and negative online expressions compared to local baselines. Lastly, I’ll describe societal sentiment responses to two events statistically attributed to human-caused climate change: the 2021 U.S. Pacific Northwest heatwave and the Western European extreme rainfall event. These results starkly reveal a fundamental aspect of human responses to emerging climatic extremes: future psychosocial impacts may far exceed those registered in the recent past, barring adaptation beyond what society has already achieved.
Kaveri Thakoor
Assistant Professor of Ophthalmic Science (in Ophthalmology), Department of Ophthalmology, Columbia University Irving Medical Center
Talk Title: Creating a Robust, Interpretable, and Portable Medical-Expert–AI Team for Eye Disease Detection
Abstract: The focus of our Artificial Intelligence for Vision Science (AI4VS) Lab is to develop AI ‘partners’ to work in tandem with clinicians to expedite eye disease detection. Our lab has 3 key goals: to robustly handle data collected from different sites/patient populations, (2) to ensure the mechanisms behind AI’s predictions are interpretable by medical experts, and (3) to create AI technology that is portable so it can reach those populations most in-need. In this lightning talk, I will give an overview of our ongoing work toward tackling these three challenges, showcasing how symbiotic expert-AI teammates may be able to achieve better disease detection accuracy and interpretability than either one alone.
Laura Kurgan
Professor of Architecture, Planning and Preservation, Graduate School of Architecture Planning and Preservation
(Moderator)
Session II: Applications of Data Science
New hardware and software are being developed to expand data science tools across a broad number of Industries.
Andrew Gelman
Higgins Professor of Statistics and Professor of Political Science, Faculty of Arts and Sciences
Talk Title: From Public Opinion to Probability Theory and Back
Abstract: What does the theory of stochastic processes have to do with public opinion? It goes like this. We learn about opinion through surveys. Most people won’t respond to a survey so we need to adjust our sample to match the population. To do this right, we need to adjust for many variables. This adjustment requires models with many parameters. To fit such a model involves exploring the needle of good fit within the haystack of possible parameter values. This exploration is performed most efficiently using programs such as Stan that use gradients and other measures of the geometry of the zone of parameter space that is consistent with data and prior information. Advanced mathematics is required to develop the fitting algorithms that use gradient information. Also the fitting should be fast: our models are all wrong, so we want to be able to fit lots of models and use graphical tools to explore the fit to data.
Stefano Fusi, Professor of Neuroscience, Vagelos College of Physicians and Surgeons; and Mingoo Seok, Associate Professor of Electrical Engineering, Columbia Engineering
Talk Title: Brain-Inspired Learning Machines
Abstract: Deep neural networks, or networks with multiple hidden layers, are revolutionizing machine learning and cognitive computing with their ability to classify images, videos and speech as rapidly and accurately as humans. The algorithms for training a deep neural network, however, are complex and require huge computational resources to tune the millions of parameters used in the training process. One way to minimize computing time is to limit each parameter’s precision from 32 bits to 1 bit. Though this reduces parameter precision, classification performance remains strong, theoretical studies show. Looking to the brain for inspiration, we are developing new methods that build on this approach.
Andreas Müller
Associate Research Scientist, Data Science Institute
Talk Title: The Rise of Open Source Software for Data Science
Abstract: The recent surge in data science applications was enabled not only by the availability of data, but also free and open software tools for processing and analyzing data. Open source projects, primarily in the R and Python programming languages, have been the backbone of most recent work in data science. The scikit-learn project in particular has become the go-to solution for machine learning algorithms in many areas. I will discuss the scikit-learn library, its scope and development, and raise questions about current model of open source tools based on volunteer labor.
John Wright
Associate Professor of Electrical Engineering, Columbia Engineering
(Moderator)
Session III: Patient Driven Healthcare
Data science is helping to diagnose, treat and prevent a range of diseases. One area of innovation has come from data collected and provided by patients themselves. We’d like to touch on some novel applications here.
Noémie Elhadad
Associate Professor of Biomedical Informatics, Vagelos College of Physicians and Surgeons
Talk Title: Citizen Endo: (Citizen + Data) Science for Endometriosis
Abstract: Ten percent of women are thought to suffer from endometriosis, a chronic condition associated with infertility and painful menstrual cycles. Despite its prevalence, we still know surprisingly little about how endometriosis develops, evolves, and how to best treat it. To learn more, my colleagues and I recently launched Citizen Endo, a crowdsourcing project that gathers data directly from endometriosis patients via a mobile phone app we developed. While most research so far has focused on physical manifestations of the disease, Citizen Endo, which now has more than 1,500 participants, aims to gain a more systemic understanding by analyzing the day-to-day experiences of women living with the condition.
Ying Cheun
Professor of Biostatistics, Mailman School of Public Health
Talk Title: Using Patient-Generated Data to Find the Best Health App
Abstract: Adaptive design is a methodology used in clinical trials that helps researchers compare different drug treatments based on interim observations of trial participants. Its interactive nature also makes it a good framework for evaluating dynamic health apps which evolve quickly over time. In this talk, I will show how adaptive design can be applied to app monitoring and recommendation within an ecosystem of health apps, and introduce SMART-AR, a framework for analyzing data in this ecosystem. I will discuss the analytics we are developing for Android mental health apps in a collaboration between Columbia and Northwestern universities.
Kenrick Cato
Assistant Professor of Nursing, School of Nursing
Talk Title: Identifying At-Risk Patients from Nursing Notes
Abstract: More than 200,000 patients die in U.S. hospitals each year from cardiac arrest, and more than 130,000 patients die of sepsis, a deadly immune response to bacterial infection. Many patient deaths could be prevented if the warning signs could be caught sooner. Our research suggests that the notes taken by nurses periodically describing their patient’s condition can provide powerful clues as to which patients need extra oversight. I will discuss a project that I am leading at Columbia, Communicating Narrative Concerns Entered by RNs (CONCERN), in partnership with several other teaching hospitals, to automate the analysis of nursing notes in patient electronic health records. Our aim is to design and evaluate clinical decision support system to identify words in nursing notes that best predict a life-threatening health problem. This project is scheduled to begin in June.
Itsik Pe’er
Associate Professor of Computer Science and
Systems Biology, Columbia Engineering
(Moderator)
Session IV: Sharing Economy
Data science is helping to remove the middle man from many industries, from finance to the media to transportation. An example is driverless cars and other changes in the transportation sector.
Costis Maglaras
David and Lyn Silfen Professor of Business and Dean, Columbia Business School
Talk Title: How Ride-hailing Platforms Optimize Performance
Abstract: Ride-hailing platforms such as Uber and Lyft match passenger demand to driver supply over a geographic network. They have at their disposal several control capabilities, such as how to prioritize passengers based on their destinations; when and which passenger requests to reject; and when and where to direct idling drivers on the network. I will discuss the impact of these type of matching and capacity optimization decisions to the network’s performance, and show that that they lead to significant improvements in most scenarios and work particularly well during the morning and evening rush hour when passenger flows across the network are most unbalanced.
Sharon Di
Assistant Professor of Civil Engineering and Engineering Mechanics, Columbia Engineering
Talk Title: A New Perspective for Shared Mobility Systems
Abstract: Ride-sharing and other shared mobility systems promise to reduce traffic, energy consumption, and the amount of time and money we spend traveling. Many of these benefits, however, remain unproven. If policymakers are to develop effective ride-sharing policies, they need to know how travelers are redistributed while using the system. In this talk I will discuss new research findings on ride-sharing that can be generalized to other shared mobility systems. This work is based on a mathematical model for predicting how travelers move across a network of roads. Early results suggest that high-occupancy toll lanes, proposed as one way to encourage carpooling and ease congestion, may be ineffective if poorly planned.
Eric L. Talley
Isidor and Seville Sulzbacher Professor of Law, Columbia Law School
Talk Title: A Machine Learning Classifier for Fiduciary Duty Waivers
Abstract: The SEC requires publicly-traded companies to periodically disclose important information to the public, but predicting how and where companies will file these disclosures is extremely difficult. I will discuss a supervised learning strategy to analyze one type of SEC disclosure filing—fiduciary duty waivers, signed by company officers and directors waiving their fiduciary duties to shareholders. Using a lawyer-coded training set of these waivers, we calibrated a predictive classifier and extended it to the entire SEC database. The results of simulated out-of-sample Monte Carlo are strong using conventional evaluation criteria. We find mildly positive effects on securities market prices, suggesting that the waiving of fiduciary duty by company officers and directors is not received as a negative signal by investors and capital market participants.
Paul Glasserman
Jack R. Anderson Professor of Business,
Columbia Business School
(Moderator)
Select Photos
Image Carousel with 9 slides
A carousel is a rotating set of images. Use the previous and next buttons to change the displayed slide
-
Slide 1: Kathleen R. McKeown
-
Slide 2: Alfred Spector
-
Slide 3: Alfred Spector
-
Slide 4: The Audience at Columbia’s Alfred Lerner Hall
-
Slide 5: 2017 Poster and Demo Session
-
Slide 6: 2017 Poster and Demo Session
-
Slide 7: 2017 Poster and Demo Session
-
Slide 8: 2017 Poster and Demo Session
-
Slide 9: 2017 Poster and Demo Session
Kathleen R. McKeown
Alfred Spector
Alfred Spector
The Audience at Columbia’s Alfred Lerner Hall
2017 Poster and Demo Session
2017 Poster and Demo Session
2017 Poster and Demo Session
2017 Poster and Demo Session
2017 Poster and Demo Session