ASA Sections on: Statistical Computing
|
[ Awards, Data expo, Video library ] [ Events, News, Newsletter ] |
Have you ever been stuck in an airport because your flight was delayed or cancelled and wondered if you could have predicted it if you'd had more data? This is your chance to find out.
The data consists of flight arrival and departure details for all commercial flights within the USA, from October 1987 to April 2008. This is a large dataset: there are nearly 120 million records in total, and takes up 1.6 gigabytes of space compressed and 12 gigabytes when uncompressed. To make sure that you're not overwhelmed by the size of the data, we've provide two brief introductions to some useful tools: linux command line tools and sqlite, a simple sql database.
The aim of the data expo is to provide a graphical summary of important features of the data set. This is intentionally vague in order to allow different entries to focus on different aspects of the data, but here are a few ideas to get you started:
You are also welcome to work with interesting subsets: you might want to compare flight patterns before and after 9/11, or between the pair of cities that you fly between most often, or all flights to and from a major airport like Chicago (ORD). Smaller subsets may also help you to match up the data to other interesting datasets.
To enter the competition you need to submit a poster to the data expo session at the 2009 JSM (more details to follow closer to the time). As well as a printed poster, you're also welcome to bring along your laptop to present interactive/animated components. After the JSM, we'll also organise a special journal issue (journal TBA) where you can submit a paper that describes your methodology in more detail.
There will be first, second, and third prizes awarded to the best posters (as judged by a panel of experts). As well as the honour and glory, the first prize consists of $1000, second prize $500, and the third prize $200. These will be awarded at the Statistical Graphics and Statistical Computing Sections Mixer at JSM 2009.
If you have any questions about the data or competition, please sign up to the competition mailing list. This will make sure that your questions can form a useful resource for other people struggling with the same problems. This mailing list will also be used to remind you of deadlines, and highlight changes to the website or data.