Project: #334 Taxi-for-all: Incentivized Taxi Actuation System for Balanced Area-wide Service Progress Report - Reporting Period Ending: Sept. 30, 2020 Principal Investigator: Carlee Joe-Wong Status: Active Start Date: July 1, 2020 End Date: June 30, 2021 Research Type: Applied Grant Type: Research Grant Program: FAST Act - Mobility National (2016 - 2022) Grant Cycle: 2020 Mobility21 UTC Progress Report (Last Updated: Oct. 5, 2020, 6:16 a.m.) % Project Completed to Date: 25 % Grant Award Expended: 20 % Match Expended & Document: 0 USDOT Requirements Accomplishments The objective of this project is to develop a method to incentivize taxis to travel to different parts of a city so that they sense data or pick up passengers in under-served areas. By incentivizing taxis to move towards locations with more passengers and higher needs, we can make the distribution of sensed data more uniform and thus useful for city operators. To accomplish this goal, we plan to develop a reinforcement learning framework to optimize the prices offered so as to make the distribution of data as even as possible. We can break down our work into three major tasks: (1) formulating the pricing problem as a reinforcement learning one, (2) designing algorithms that efficiently learn the optimal prices within this framework, and (3) evaluating this algorithm compared to baselines on taxi and Roadbotics datasets, as well as in a small Shenzhen deployment. Since the start of the project in July, we have begun to formulate the incentives problem in a reinforcement learning framework. We have decided on a preliminary definition of the states, actions, and reward function. However, a major challenge is that our current reward is not separable over time (as is usually assumed in reinforcement learning problems), and our current definition of the state is quite large, which may lead to scalability problems. We are currently working on designing reinforcement learning variants to address these challenges. We have also begun to build a simulation platform for taxi movement around a city as a function of the incentives offered to taxi drivers, which will allow us to test our formulation and algorithms. As part of the simulator construction, we have implemented existing reinforcement learning techniques that can take as input generic state and reward variables and attempt to learn the optimal actions. Currently, the simulator integrates data on taxi movement from Beijing. Training and professional development opportunities have been provided in the form of taxi datasets that have been used for projects in a data analysis course at Stanford. We plan to use these datasets in other courses in later semesters as well. One workshop paper on our simulation platform was presented in September. We plan to continue disseminating our results by publishing papers on our frameworks and algorithms over the course of the project. Our goal for the next six month reporting period is to continue refining our pricing algorithms and testing them on our simulator. Impacts Our results to date have increased scientific knowledge by identifying challenges in applying reinforcement learning to problems with a large number of spatial state variables, in particular scalability challenges as the size of the state space grows. We believe that this challenge will arise in general spatiotemporal learning scenarios beyond taxi incentivization, where an action variable must be optimized at each location over a large geographical space. The algorithms we are currently developing may thus be useful for a more general class of reinforcement learning problems. Our current simulation framework will allow researchers to prototype and test ideas for incentivizing taxi movement around a city, including our own planned work in this project. We are working on releasing an open-source version of the simulator to supplement our existing workshop publication, so that it may be used by other researchers and companies. As we finalize our pricing algorithms and progress further into the project, we will investigate more opportunities for technology transfer. In particular, as stated in the proposal we plan to explore collaborations with Roadbotics, which uses a fleet of vehicles to collect sensing data in Pittsburgh; and with taxi companies in Shenzhen, China. Other We report two major outputs for this reporting period. The first is the development of a software simulator for taxi drivers’ reactions to incentives offered to them. The second is the development of a reinforcement learning framework to optimize these incentives. Our simulator, which is the basis for our workshop paper, includes modules that admit customized models for how taxi drivers react to the incentives given, e.g., some drivers may simply ignore the incentives, while others may attempt to optimize their driving destinations so as to maximize the incentives received. The simulator utilizes traces of taxi activity from Beijing, China; and models the effect of taxi driver movement on congestion throughout a city. We are working on releasing an open-source version of the simulator that would allow other researchers to use our work in their own research on incentivization in transportation networks. Our reinforcement learning framework, as described above, specifies the prices chosen throughout the city as actions taken so as to maximize a specified reward function, which we define as the degree to which the taxi distribution matches a specified target distribution. These techniques will allow sensing or taxi operators to ensure that taxis spread themselves around a city according to sensing or passenger needs. Outcomes New Partners None to report. Issues No significant changes to report. COVID-19 may limit our ability to deploy our algorithms in taxis, but as the deployment is not planned until the last few months of the project, we have not yet made any definitive changes to our plans.