Parallelize code3.py with Pool+map. 

1. rewrite code3.py with build in map function.

2. Parallelize it with Pool+map  (see code2.py as an example).

3. Make sure that matrix multipication use only one thread per  
   process (use "mkl.set_num_threads(1)" in the code and "export
   OMP_NUM_THREADS=1" in the batch file).  

4. Compare perforamance of code3.py with your new code.

5. We've made sure that we use only one thread per process. Now, try to
   set 12 threads per process (set "mkl.set_num_threads(12)" in the code
   and OMP_NUM_THREADS=12 in the batch file) and see how performance will
   degrade.
