Posts

Showing posts with the label json

JSON Performance

I have some code that processes tweets, about 5 million a day, in realtime. They are currently stored in mongodb and also posted on various celery/rabbitmq work queues. The average message size is 5524, so encoding and decoding these messages is an issue. Using the following test code below. Standard Tweet message Encode/Decode with python built-in json package. message size Test Msg Size De-serialize Obj Size Serialize cjosn, bson, ujson Storage Storoge Cost Empty object 41 14.790, 37.565, 0.970 54 6.341, 41.856, 1.249 2050Mb 0.21 {} Empty list 41 15.069, 38.021, 1.005 54 6.675, 41.475, 1.400 2050Mb 0.21 [] Object of objects 843 107.750, 145.440, 25.525 3226 63.555, 828.235, 28.051 42150Mb 4.21 List of lists 563 58.805, 81.950, 16.960 104 43.426, 815.965, 18.311 28150Mb 2.81 Object with only tweet id 93 25.030, 53.360, 2.280 422 23.570, 83.445, 3.295 4650Mb 0.47 Full tweet message 4386 697.221, 867.780, 188.560 12606 360.290, 5847.335, 201.610 219300Mb 21...

Serializion Performance

Last week  I stuck my head out  in a meeting and declared that XML is verbose and slow to parse and that we should move to something like Google's protocols buffers,  or something readable such as json or YAML, which are  easier to parse etc etc etc! Well is this really true ? The statement seems logical considering how verbose XML can be. Still, after the meeting, some questions stayed in my mind. So I thought I would do some tests. I used  a FIX Globex (CME) swap trade confirmation message to test my theory. Size from Python to Python json cjson 2332 0.222238063812 0.0943419933319 pickle cPickle 1778 0.233518123627 0.128826141357 XML cElementTree 2083 0.407706975937 2.77832698822 json simplejson 2332 3.37723612785 5.11316084862 So this simple test shows that using XML with cElementTree parser  is not so slow, cjson wins in speed and the conclusion must be: Your performance will ultimately depend on your data and the quality of the l...