{"id":1381,"date":"2011-07-25T19:11:50","date_gmt":"2011-07-25T23:11:50","guid":{"rendered":"https:\/\/www.bu.edu\/pasi\/materials\/post-pasi-training-week-2\/"},"modified":"2011-08-19T12:06:41","modified_gmt":"2011-08-19T16:06:41","slug":"post-pasi-training-week-2","status":"publish","type":"page","link":"https:\/\/www.bu.edu\/pasi\/materials\/post-pasi-training\/post-pasi-training-week-2\/","title":{"rendered":"Post-PASI training: Week 2"},"content":{"rendered":"<h2>Syllabus<\/h2>\n<p><strong>August 15th:<\/strong><\/p>\n<ul>\n<li>Holiday<\/li>\n<\/ul>\n<p><strong>Lab 3 (August 16th):<\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/Lab3.pdf\">Lab3 &#8211; Slides<\/a><\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/FD_2D_shared.cu_1.zip\">FD_2D_shared.cu<\/a><br \/>\n<\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/FD_2D_shared_ghost.cu_.zip\">FD_2D_shared_ghost.cu<\/a><br \/>\n<\/strong><\/p>\n<ul>\n<li>Using shared memory as cache\n<ul>\n<li>Implement 2D explicit heat transfer with shared memory<\/li>\n<\/ul>\n<\/li>\n<li>Comparison of each implementation: timings vs programming effort<\/li>\n<\/ul>\n<ul>\n<li><strong>References: <\/strong>\n<ul>\n<li>Kirk, D. and Hwu, W. Programming Massively Parallel Processors.(<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter4-CudaMemoryModel.pdf\">Ch. 4<\/a>,\u00a0<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter5-CudaPerformance.pdf\">Ch. 5<\/a>)<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>Class 4 (August 17th):<\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/Lecture4.pdf\">Lecture 4 &#8211; Slides<\/a><br \/>\n<\/strong><\/p>\n<ul>\n<li>Control flow\n<ul>\n<li>Warp divergence<\/li>\n<\/ul>\n<\/li>\n<li>Memory coalescing<\/li>\n<li>Latency hiding<\/li>\n<li>Occupancy<\/li>\n<li>Measuring effective performance<\/li>\n<\/ul>\n<p><strong>Lab 4 (August 18th):<\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/Lab4.pdf\">Lab4 &#8211; Slides<\/a><\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/lab4_files.zip\">lab4_files<\/a><\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/AAt_tiled.cu_.zip\">AAt_tiled.cu<\/a><br \/>\n<\/strong><\/p>\n<p><strong> <\/strong><\/p>\n<ul>\n<li>Implement an efficient A A_transpose multiplication<\/li>\n<li><strong>References: <\/strong>\n<ul>\n<li>Kirk, D. and Hwu, W. Programming Massively Parallel Processors. (<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter4-CudaMemoryModel.pdf\">Ch. 4<\/a>,\u00a0<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter5-CudaPerformance.pdf\">Ch. 5<\/a>)<\/li>\n<li>NVIDIA Advanced CUDA Webminar. Memory Optimizations (<a href=\"http:\/\/developer.nvidia.com\/gpu-computing-webinars\">http:\/\/developer.nvidia.com\/gpu-computing-webinars<\/a>)<\/li>\n<li><a href=\"http:\/\/developer.download.nvidia.com\/compute\/cuda\/4_0_rc2\/toolkit\/docs\/CUDA_C_Best_Practices_Guide.pdf\">CUDA C Best Practices Guide. Ch. 2-6.<\/a><\/li>\n<li><a href=\"http:\/\/www.cs.colostate.edu\/~cs675\/MatrixTranspose.pdf\">Ruetsch, G. and Micikevicius, P. Optimizing Matrix Transpose in CUDA.<\/a><\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>Class 5 (August 19th):<\/strong><\/p>\n<p><strong><a href=\"\/pasi\/files\/2011\/07\/Lecture5.pdf\">Lecture 5 &#8211; Slides<\/a><br \/>\n<\/strong><\/p>\n<ul>\n<li>Further optimization techniques:\n<ul>\n<li>Data prefetching<\/li>\n<li>Instruction optimization<\/li>\n<li>Loop unrolling<\/li>\n<\/ul>\n<\/li>\n<li>Thread and block heuristics<\/li>\n<li>Example: optimizing a parallel reduction<\/li>\n<\/ul>\n<ul>\n<li><strong>References: <\/strong>\n<ul>\n<li>Kirk, D. and Hwu, W. Programming Massively Parallel Processors. (<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter4-CudaMemoryModel.pdf\">Ch. 4<\/a>,\u00a0<a href=\"http:\/\/courses.engr.illinois.edu\/ece498\/al\/textbook\/Chapter5-CudaPerformance.pdf\">Ch. 5<\/a>)<\/li>\n<li><a href=\"http:\/\/developer.download.nvidia.com\/compute\/cuda\/1_1\/Website\/projects\/reduction\/doc\/reduction.pdf\">Mark Harris, NVIDIA. Optimizing Parallel Reduction in CUDA<\/a><\/li>\n<\/ul>\n<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Syllabus August 15th: Holiday Lab 3 (August 16th): Lab3 &#8211; Slides FD_2D_shared.cu FD_2D_shared_ghost.cu Using shared memory as cache Implement 2D explicit heat transfer with shared memory Comparison of each implementation: timings vs programming effort References: Kirk, D. and Hwu, W. Programming Massively Parallel Processors.(Ch. 4,\u00a0Ch. 5) Class 4 (August 17th): Lecture 4 &#8211; Slides Control [&hellip;]<\/p>\n","protected":false},"author":3344,"featured_media":0,"parent":1297,"menu_order":2,"comment_status":"closed","ping_status":"closed","template":"","meta":[],"_links":{"self":[{"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/pages\/1381"}],"collection":[{"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/users\/3344"}],"replies":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/comments?post=1381"}],"version-history":[{"count":19,"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/pages\/1381\/revisions"}],"predecessor-version":[{"id":1532,"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/pages\/1381\/revisions\/1532"}],"up":[{"embeddable":true,"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/pages\/1297"}],"wp:attachment":[{"href":"https:\/\/www.bu.edu\/pasi\/wp-json\/wp\/v2\/media?parent=1381"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}