Project: #334 Taxi-for-all: Incentivized Taxi Actuation System for Balanced Area-wide Service Progress Report - Reporting Period Ending: March 31, 2021 Principal Investigator: Carlee Joe-Wong Status: Active Start Date: July 1, 2020 End Date: June 30, 2021 Research Type: Applied Grant Type: Research Grant Program: FAST Act - Mobility National (2016 - 2022) Grant Cycle: 2020 Mobility21 UTC Progress Report (Last Updated: March 30, 2021, 5:05 p.m.) % Project Completed to Date: 75 % Grant Award Expended: 0 % Match Expended & Document: 0 USDOT Requirements Accomplishments The objective of this project is to develop a method to incentivize taxis to travel to different parts of a city so that they sense data or pick up passengers in under-served areas. By incentivizing taxis to move towards locations with more passengers and higher needs, we can make the distribution of sensed data more uniform and thus useful for city operators. To accomplish this goal, we plan to develop a reinforcement learning framework to optimize the prices offered so as to make the distribution of data as even as possible. We can break down our work into three major tasks: (1) formulating the pricing problem as a reinforcement learning one, (2) designing algorithms that efficiently learn the optimal prices within this framework, and (3) evaluating this algorithm compared to baselines on taxi and Roadbotics datasets, as well as in a small Shenzhen deployment. Since the start of the project in July, we have begun to formulate the incentives problem in a reinforcement learning framework. We have decided on a preliminary definition of the states, actions, and reward function, and have begun running experiments with these algorithms. A key challenge, however, has been the scalability of our algorithms to realistic time durations and city sizes. We have begun investigating the use of clustering algorithms to reduce the size of our state space and learn policies with a manageable number of free variables. We have also built a simulation platform for taxi movement around a city as a function of the incentives offered to taxi drivers, which will allow us to test our formulation and algorithms. As part of the simulator construction, we have implemented existing reinforcement learning techniques that can take as input generic state and reward variables and attempt to learn the optimal actions. Currently, the simulator integrates data on taxi movement from Beijing. Training and professional development opportunities have been provided in the form of taxi datasets that have been used for projects in a data analysis course at Stanford. We plan to use these datasets in other courses in later semesters as well. One workshop paper on our simulation platform was presented in September. We plan to continue disseminating our results by publishing papers on our frameworks and algorithms over the course of the project. In particular, we are working on a journal paper submission that outlines our reinforcement learning framework ideas. Our goal for the next six month reporting period is to continue refining our pricing algorithms and testing them on our simulator. We will further investigate a variation on our reinforcement learning problem that examines how taxis should act given their locations and mobility. For example, if taxis can display digital advertisements, different advertisements may be more effective at different locations. It is difficult to determine their effectiveness in advance, making the problem of “which advertisement to display” amenable to reinforcement learning solutions. Impacts Our results to date have increased scientific knowledge by identifying challenges in applying reinforcement learning to problems with a large number of spatial state variables, in particular scalability challenges as the size of the state space grows. We believe that this challenge will arise in general spatiotemporal learning scenarios beyond taxi incentivization, where an action variable must be optimized at each location over a large geographical space. The algorithms we are currently developing, which intelligently cluster state variables to learn a reduced policy that is nonetheless almost optimal, may thus be useful for a more general class of reinforcement learning problems. Our current simulation framework will allow researchers to prototype and test ideas for incentivizing taxi movement around a city, including our own planned work in this project. We are working on releasing an open-source version of the simulator to supplement our existing workshop publication, so that it may be used by other researchers and companies. As we finalize our pricing algorithms and progress further into the project, we will investigate more opportunities for technology transfer. In particular, as stated in the proposal we plan to explore collaborations with Roadbotics, which uses a fleet of vehicles to collect sensing data in Pittsburgh; and with taxi companies in Shenzhen, China. We are additionally discussing the possibility of obtaining taxi movement data from DriveSally, a company that rents taxi vehicles to drivers. Our algorithms can help DriveSally recommend where their drivers should go so as to maximize their revenue. Other We report two main outcomes for this reporting period. The first is the development of a reinforcement learning framework to optimize the incentives offered to taxi drivers for moving to different locations around a city. As described above, this framework specifies the prices chosen throughout the city as actions taken so as to maximize a specified reward function, which we define as the degree to which the taxi distribution matches a specified target distribution. These techniques will allow sensing or taxi operators to ensure that taxis spread themselves around a city according to sensing or passenger needs. Our algorithms include new techniques for efficiently learning near-optimal policies with a large possible state space, by intelligently clustering some states together and condensing their policies. Our second main outcome for this reporting period is an initial formulation of the related problem of how taxis should act once they have traveled to different locations. While it is fairly straightforward to say that taxis should pick up any available passengers, other potential actions may not be as obvious. Taxis displaying digital advertisements, for example, or equipped with multiple sensors, may need to decide which advertisements to display or which data to upload. Our decision framework allows them to learn over time which actions to take at which locations. Outcomes New Partners We have begun discussions with DriveSally, a company that rents vehicles to taxi and rideshare drivers. They also have a side business displaying digital advertising on their vehicles. DriveSally is interested in giving us access to some of their taxi mobility data, which we will be able to use to simulate taxi movements in response to different driver incentives. In particular, they would like to use our algorithms to incentivize their drivers to locations where there are more potential passengers and they can earn revenue by displaying effective digital advertisements. Issues COVID-19 may limit our ability to deploy our algorithms in taxis in Shenzhen, as was originally planned. To compensate for the loss of this trial, we are talking with Drive Sally, a company that rents cars to be used for ridesharing or taxis, about obtaining their datasets on vehicle movement. These datasets will enable us to better simulate taxi movement around a city, and DriveSally is interested in using our algorithms to optimize their vehicle movements. Our reinforcement learning algorithms will allow Drive Sally to incentivize drivers to head towards the locations where advertisements will be the most effective.