You guys should check out our high performance GPU graph processing library, Gunrock: http://gunrock.github.io/gunrock/ We are working on multi-GPU distributed version now.
Interesting -- though a quick scan of the evaluation data sets suggests that none of them are as large as the Twitter one (1.1B edges) or the uk-2007-05 one (3B edges) that we (and other distributed graph processing systems) use. Presumably this is due to memory limitations on the GPU?