Outside the Classroom: Data Science Corps, Q & A with DSAN Director Dr. Purna Gamage
What inspired you to join the Data Science Corps program this year?

The opportunity came through a broader Georgetown collaboration around the NSF Data Science Corps program. I was excited to be part of this NSF-funded initiative with colleagues in leadership roles across campus, including Dr. Mahlet Tadesse, Chair of the Department of Mathematics and Statistics; Dr. Lisa Singh, who was Director of the Massive Data Institute at the time; Dr. Britt Qiwei Hu from DSAN; and myself, as Director of the Data Science and Analytics program. What made the project especially meaningful was that it brought together different parts of Georgetown — Mathematics and Statistics, Computer Science, MDI, and DSAN — around a shared goal of supporting applied data science research, teaching, and student-centered learning.
The program creates a space where undergraduate students and K–12 teachers can work collaboratively on real data science research projects while also thinking about how data science concepts can be translated into curriculum and classroom practice.
As one of the faculty leads for a research group this summer, I saw this as a valuable opportunity to mentor students and teachers through a hands-on project connected to timely questions about AI, labor-market change, and data-driven decision making. The program is supported by the U.S. National Science Foundation under award No. DRL-2436779, as part of the NSF Data Science Corps program.

The program brings together high school teachers and undergraduate researchers. What have you seen happen when those two groups actually sit down together?
What I have really enjoyed seeing is how much both groups learn from each other. The undergraduate researchers usually come in with more recent coding and technical experience, while the teachers bring a very strong understanding of how to explain ideas clearly. When they sit down together, it becomes a very collaborative environment.
Your DS Corps project examines how AI adoption is reflected in labor-market demand, job openings, and skill requirements. It brings together complex data sources and methods, including job postings, Google Trends, economic indicators modeling with time-series forecasting, Deep Learning Modeling and NLP analysis. How do you help students and teachers engage with these advanced data science methods in a way that is accessible, but still rigorous and not oversimplified?

I think the key is to start with the research question, not the method. We first help students and teachers understand the real-world question: how is AI adoption showing up in the labor market, and what signals might help us study that? Once they understand the question, the methods become less intimidating because each method has a clear purpose.
For example, job postings tell us what is happening in the labor market, Google Trends gives us a sense of public interest, economic indicators give us broader context, and time-series models help us study how these patterns change over time. Deep learning and NLP are more advanced methods, but we introduce them as tools for specific tasks rather than as abstract concepts.
We also spend time connecting the project to the existing literature on AI exposure, labor-market change, job-skill demand, and forecasting. This helps participants see that the project is part of a larger research conversation, not just a coding exercise. It also gives them a stronger reason for why we use certain data sources and methods.
We then break the project into smaller steps. Students and teachers work through data cleaning, visualization, exploratory analysis, modeling, and interpretation one piece at a time. We use hands-on coding examples, weekly discussions, and presentation practice so they can build confidence gradually.
At the same time, we do not want to oversimplify the work. We still talk about model assumptions, prediction error, uncertainty, and why one model may perform better than another. The goal is not just to run models, but to help participants understand what the results mean and what limitations come with them.
So for me, accessibility means giving enough structure and support so students and teachers can engage with advanced methods, while rigor means making sure they still understand the data, the modeling choices, and the interpretation behind the results.
How has the DSAN program contributed to the DS Corps experience, and how has DS Corps benefited from DSAN’s experience in applied data science education, especially in supporting both students and teachers?
DSAN played a major role in the DS Corps experience because the program needed both technical depth and applied, hands-on data science training. Dr. Britt and I were both leading contributors to the program, and since we are both from DSAN, we were able to bring in the program’s strong experience with applied data science education, project-based learning, and working with real-world data.
For this summer, we also shared the DSAN Bootcamp coding materials with the DS Corps participants to help them build their coding foundation. That was very helpful because many of the students and teachers were learning new tools while also working on a research project. The bootcamp materials gave them extra support with coding, data cleaning, visualization, and analysis.
For this project, I led lectures and guidance on time-series analysis and forecasting, while Dr. James Hickman supported the group with deep learning methods. These topics were central to the research because the students and teachers were not only learning data science concepts, but also applying them directly to a real research question about AI adoption and labor-market trends. The DSAN experience was especially important here because our program regularly teaches students how to move from messy real-world data to modeling, interpretation, and communication of results.
We worked closely with the students and teachers throughout the summer, helping them develop research questions, work through coding challenges, understand the models, interpret the results, and prepare for weekly presentations. The project was very hands-on, and I think that made the learning experience much stronger.
A major part of the program’s success was also due to our RA, Nikhil Patla, who is a second-year DSAN student. He was an amazing help throughout the summer. He spent many hours supporting the research and coding work, helping students troubleshoot, practicing presentations with them every week, and staying in constant communication with me and the team. His support made a big difference in keeping the project moving and helping the students and teachers feel supported throughout the process.
What would you want a DSAN student reading this to take away from knowing this program exists?
I would want DSAN students to see that the skills they are learning in the program can have an impact beyond the classroom. DS Corps is a good example of how applied data science can support real research, help teachers bring data science into their classrooms, and give students the chance to work on meaningful problems with real data.
I also hope they see that DSAN students have a lot to contribute. Nikhil’s work as the RA is a great example of this. He was not only helping with coding and analysis, but also mentoring participants, supporting presentations, and helping the project move forward. That kind of contribution shows how valuable DSAN training can be when it is applied in a real collaborative setting.
So the main takeaway is that DSAN is not just about learning models or coding skills. It is about using those skills to support people, answer real questions, and communicate results in a way that others can understand and use.