Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 1 of 15 EXHIBIT 2 Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 2 of 15 USOO8648867B2 (12) United States Patent (10) Patent No.: US 8,648,867 B2 Gorchetchnikov et al. (45) Date of Patent: Feb. 11, 2014 (54) GRAPHIC PROCESSORBASED (56) References Cited ACCELERATOR SYSTEMAND METHOD U.S. PATENT DOCUMENTS (75) Inventors: Anatoli Gorchetchnikov, Belmont, MA 5,388,206 A * 2/1995 Poulton et al. ................ 345,505 (US); Heather Marie Ames, South 2005, 0166042 A1* 7, 2005 Evans ............ T13,150 Boston, MA (US); Massimiliano 2007/0052713 A1* 3/2007 Chung et al. ... 34.5/5O1 Versace, South Boston, MA (US); 2007/0279429 A1* 12/2007 Ganzer ......................... 345,582 Fabrizio Santini, Jamaica Plain, MA (US) * cited by examiner (73) Assignee: Neurala LLC, Boston, MA (US) Primary Examiner — Maurice L. McDowell, Jr. (*) Notice: Subject to any disclaimer, the term of this (57) ABSTRACT patent is extended or adjusted under 35 An accelerator system is implemented on an expansion card U.S.C. 154(b) by 1030 days. comprising a printed circuit board having (a) one or more graphics processing units (GPU), (b) two or more associated (21) Appl. No.: 11/860,254 memory banks (logically or physically partitioned), (c) a (22) Filed: Sep. 24, 2007 specialized controller, and (d) a local bus providing signal coupling compatible with the PCI industry standards (this (65) Prior Publication Data includes but is not limited to PCI-Express, PCI-X, USB 2.0, or functionally similar technologies). The controller handles US 2008/O 117220 A1 May 22, 2008 most of the primitive operations needed to set up and control Related U.S. Application Data GPU computation. As a result, the computer's central pro cessing unit (CPU) is freed from this function and is dedicated (60) Provisional application No. 60/826,892, filed on Sep. to other tasks. In this case a few controls (simulation start and 25, 2006. stop signals from the CPU and the simulation completion signal back to CPU), GPU programs and input/output data are (51) Int. C. the information exchanged between CPU and the expansion G06F 5/00 (2006.01) card. Moreover, since on every time step of the simulation the (52) U.S. C. results from the previous time step are used but not changed, USPC ........................................... 345/501; 34.5/503 the results are preferably transferred back to CPU in parallel (58) Field of Classification Search with the computation. USPC .................................................. 345/501,503 See application file for complete search history. 19 Claims, 5 Drawing Sheets 140 200 8O 250 -4 SHADERE.K. MEMORY 50 O EXRE - icon TROLLER w NEMORY C - 2 Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 3 of 15 U.S. Patent Feb. 11, 2014 Sheet 1 of 5 US 8,648,867 B2 7. $3. s FG, Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 4 of 15 U.S. Patent Feb. 11, 2014 Sheet 2 of 5 US 8,648,867 B2 g i. : i : r s al e3esiasti Od s 3 s Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 5 of 15 U.S. Patent Feb. 11, 2014 Sheet 3 of 5 US 8,648,867 B2 300 CPU 120 Start 304 Expansion Cardiso : iaia fict iss: Risitio: &::::::::8::::::::: Sires 301 Se::303 320 Disk:O Graphic User Controller Interface initialization initialization Initialization 325 User input Parser External input Texture textures from RAM to interaction Generator texture memory bank 326 : Population Parser Population shader Shader Generato binaries from RAM to and Compiler shader memory bank Simulation Initialization 330 Progress Input Parser monitor Texture Generator Computation (see Fig. 4) Data output to Disk Output Data Data ACCumulation issils in RAM Last iteration 318 : Yes 350 3 Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 6 of 15 U.S. Patent Feb. 11, 2014 Sheet 4 of 5 US 8,648,867 B2 Expansion Card 180 Simulation start iDesi: {34}{ut : {iciplisiii) 2. iaia iiii: Si:Sires: Siii:Sire:: Silisirai? 402 403 404 Input textures from texture memory bank to GPU Shaders from shader memory bank to GPU ... ........... Shader execution Output texture upload to texture memory bank New external input Data textures from RAM told ----------- Wait for swap Swap input Wait for swap of input/output and output of input/output texture pointers texture pointers texture pointers 475 Input textures from NO Last texture memory iteration? bank to RAM Yes iteration? Yes 490 Wait for all three streams of execution to finish 499 EIG 4 Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 7 of 15 U.S. Patent Feb. 11, 2014 Sheet 5 of 5 US 8,648,867 B2 (p) ··*,?)&xaddO2OInputDependency(m gate1, O., O., 0.004, O., 0, 0); Interaction Stream, Data Output Stream, and Computational u->addFullDependency(m gate2, population()); Stream. Through TSimulator:timestep(), TSimulator:out 40 fileInterval ( ), and TSimulator:outmode(), the application public: TCablePopRCF(): TPopulation(“compCPU RCF, w, h, true) can set the time step of the simulation, the time step of disk { }; output, and the mode of the disk output. The external input ~TCablePopRCF() {if(m equation) delete m equation; if(m gate1) delete m gate1; pattern should be packed into a TPattern object and bound to the simulation object through TSimulator:resetInputs( ) 45 bool if(m gate2) delete m gate2:}; fillElements(TSimulator sim); method. TSimulator::simLength( ) sets the length of the }: simulation. bool TCablePopRCF::fillElements(TSimulatior sim) The second step is to create at least one population of { equation = new TEq RCF (this, m compet, m persist); equations (TPopulation object). Population holds one equa mcreateCatingStructure(); tion object TEquation. This object contains only a formula 50 for(size ti = 0; i < x.Size(); ++i) and does not hold element-specific data, so all elements of the for(size tj = 0; j < ySize(); ++) population can share single TEquation. { The TEquation object is converted to a GPU program TElement u = new TCPUElement(this, m equation, i,j): sim->addUnit(u); before execution. GPU programs have to be executed within createUnitStructure(u); a graphical context, which is stream specific. TSimulator 55 return true: creates this context within a Computational Stream, therefore all programs and data arrays that are necessary for computa int tion have to be initialized within Computational Stream. Con main() structor of TPopulation is called from User Interaction //{ Input pattern generation (309 in FIG. 3) Stream, so no GPU-related objects can be initialized in this 60 uint32 t pat = new uint32 twh; COnStructOr. TRandom-float randGen (O); TPopulation::fillElements( ) is a virtual method designed for(uint32 ti = 0; i < wh; ++i) to overcome this difficulty. It is called from within the Com pati = randGen.random (); putational Stream after TSimulator::networkCreate( ) is TPattern p = new TPattern (pat, w, h); if Setting up the simulation called in the User Interaction Stream. A user has to override 65 TSimulator cableSim = new TSimulator(“data'); fi(308 and 320 in TPopulation::fillElements( ) to create TEquation and other FIG. 3) computation related objects both element independent and Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 13 of 15 US 8,648,867 B2 11 12 -continued need to perform the computations and initiates the upload 435 of them onto the GPU 240. The GPU 240 can communicate cableSim->timestep(0.05); fi(320 in FIG. 3) cableSim->resetInputs(p); (325 in FIG. 3) directly with the texture memory bank 250 to upload the cableSim->OutfileInterval(0.1); (308 in FIG. 3) appropriate texture to perform the computations. The control cableSim->Outmode(SANNDRA::timefunc); (308 in FIG. 3) cableSim->simLength(60.0); (320 in FIG. 3) ler 220 also pulls the first shader (known by the stored order) if Preparing the population from the shader memory bank 210 and uploads 450 it onto the TPopulation* cablePop = new TCablePopRCF(); //(310 in FIG. 3) GPU 240. cableSim->networkCreate(); //(326 in FIG. 3) The GPU 240 executes the following operations in this uint16 t user = 1; order: performs the computation (execution of the shader) while(user) 10 470; tells the controller 220 that it is done with the computa if(cableSim->simulationStart(true, 1)) (311 in FIG. 3) tions for the current shader; and after all shaders for this exit(1): particular equation are executed sends 480 the output textures std::cout-“Repeat?\n: //(305 in FIG. 3) std::cin>user; //(305 in FIG. 3) to the output portion of the texture memory bank 250. This if(user == 1) 15 cycle continues through all of the equations based on the cableSim->networkReset(); //(305 in FIG. 3) branching step 482. if cableSim) An example shader that performs fourth order Runge delete cableSim; Also deletes cablePop and its internals Kutta numerical integration is shown in Listing 2 using GLSL exit(0); notation; Listing 1. FIG. 4 is a detailed flow diagram illustrating a part of an uniform sampler2DRect Variable; uniform float integration step; exemplary implementation of the bottom level system and float halfstep = integration step*0.5; method performed during the computation on the GPU accel 25 float fl. 6 step = integration stepf 6.0: erator of the expansion card 180 and is a more detailed view vec4 output = texture2DRect(Variable, gl TexCoord O.st); if define equation() here of the computational box 330 in FIG. 3. FIG. 4 is a represen vec4 rungekutta4(vec4 x) tation of one of several ways in which a system and method for processing numerical techniques can be implemented. const vecA k1 = equation(x); const vec4 k2 = equation(x + halfstep*k1); With systems of equations that have complex interdepen 30 const vec4 k3 = equation(x + halfstep*k2); dencies it is likely that the variable in Some equation from a const vecA k4 = equation(x + integration Step*k3); previous time step has to be used by some other equation after return fl. 6step*(k1 + 2.0* (k2 + k3) + k4); the new values of this variable are already computed for new time step. To avoid data confusion, the new values of variables void main (void) { should be rendered in a separate texture. After the time step is 35 output += rungekutta4(output); completed for all equations, these new values should be cop gl FragColor = output; ied over old values so that they are used as input during the Listing 2. next time step. Copying textures is an expensive operation, computationally, but since the textures are referred to by texture IDs (pointers), Swapping these pointers for input and 40 The shader in Listing 2 can be executed on conventional output textures after each time step achieves the same resultat video card. Using the controller 220 this code can be further a much lesser cost. optimized, however. Since the integration step does not In the hardware solution suggested herein, ID Swapping is change during the simulation, the step itself as well as the equivalent to Swapping the base memory address for two halfstep and /6 of the step can be computed once per simula partitions of the texture memory bank 250. They are swapped 45 tion, and updated in all shaders by a shader update procedures 485 during synchronization (485, 430, and 455) so that data transfer 445 and the computation 435-487 proceeds immedi 310,326 discussed above. ately and in parallel with data transfer as shown in FIG. 4. A computed theofmain After all the equations in the computational cycle are hardware solution allows this parallelism through access of 220 can switch 485execution the substream 403 on the controller reference pointers of the input and the controller 220 to the onboard texture memory bank 250. 50 output portions of the texture memory bank 250. The main computation and data exchange are executed by The two other substreams of execution on the controller the controller 220. It runs three parallel substreams of execu tion: Computational Substream 403, Data Output Substream 220 are waiting (blocks 430 and 455, respectively) for this 402, and Data Input Substream 404. These streams are syn switch to begin their execution. The Data Input Substream chronized with each other during the swap of pointers 485 to 55 404 is controlling 440 the input of additional data from the the input and output texture memory partitions of the texture CPU 120. This is necessary in cases where the simulation is memory bank 250 and the check for the last iteration 487. monitoring the changing input, for example input from a Algorithmically, these two operations are a single atomic video camera or other recording device in the real time. This operation, but the block diagram shows them as two separate substream uploads new external input from the CPU 120 to blocks for clarity. 60 the texture memory bank 250 so it can be used by the main The Computational Substream 403 performs a computa computational Substream 403 on the next computational step tional cycle including a sequential execution of all shaders and waits for the next iteration 475. The Data Output Sub that were stored in the shader memory bank 210 using the stream 445 controls the output of simulation results to the appropriate input and output textures. To begin the simulation CPU 120 if requested by the user. This substream uploads the the controller 220 initializes three execution substreams 403, 65 results of the previous step to the main RAM 130 so that the 402, and 404. On every simulation step, the Computational CPU 120 can save them on disk 140 or show them on the Substream 403 determines which textures the GPU 240 will results display 313 and waits for the next iteration 460. Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 14 of 15 US 8,648,867 B2 13 14 Since the Computational Substream 403 determines the While this invention has been particularly shown and timing of input 440 and output 445 data transfers, these data described with references to preferred embodiments thereof, transfers are driven by the controller 220. To further reduce it will be understood by those skilled in the art that various the data transfer overhead (and disk 140 overhead also) the changes in form and details may be made therein without controller 220 initiates transfer only after selected computa departing from the scope of the invention encompassed by the tional steps. For example, if the experimental data that is appended claims. simulated was recorded every 10 milliseconds (msec) and the simulation for better precision was computed every 1 mSec, What is claimed is: then only every tenth result has to be transferred to match the 1. A computer system for performing a numerical simula experimental frequency. 10 tion over a plurality of computational cycles including at least This solution stores two copies of output data, one in the a first computational cycle and a second computational cycle, expansion card texture memory bank 250 and another in the the computer system comprising: system RAM 130. The copy in the system RAM 130 is a central processing unit; accessed twice: for disk I/O and screen visualization 313. An a main memory, operably coupled to the central processing alternative solution would be to provide CPU 120 with a 15 unit, to store input data to be accessed by the central direct read access to the onboard texture memory bank 250 by processing unit in performing the numerical simulation; mapping the memory of the hardware onto a global memory a video system, operably coupled to the central processing space. The alternative solution will double the communica unit, to drive a video monitor to display an indication of tion through the local bus 190. Since the goal discussed herein the numerical simulation in response to the computer is reducing the information transfer through the local bus 190, system performing the numerical simulation; the former solution is favored. an accelerator, operably coupled to the central processing The main stream substream 403 determines if this is the last unit, to receive at least a portion of the input data from iteration 487. If it is the last iteration, the controller 220 waits the central processing unit and to provide first output for the all of the execution substreams to finish 490 and then data generated during the first computational cycle to the returns the control to the CPU 120, otherwise it begins the 25 central processing unit after a conclusion of the first next computational cycle. computational cycle, the accelerator comprising: This repeats through all of the computational cycles of the at least one graphics processing unit to generate second simulation. output data, during the second computational cycle, Conclusion by performing at least one calculation on the first This GPU accelerator system offers the following potential 30 output data; and advantages: an accelerator memory, operably coupled to the at least 1. Limited computations on the CPU 120. The CPU 120 is one graphic processing unit, the accelerator memory only used for user input, sending information to the controller comprising: 220, receiving output after each computational cycle (or less a first partition, referenced by a first pointer, to store frequently as defined by the user), writing this output to disk 35 the first output data during the second computa 140, and displaying this output on the monitor 170. This frees tional cycle; and the CPU 120 to execute other applications and allows the a second partition, referenced by a second pointer, to expansion card to run at its full capacity without being slowed store the second output data generated during the down by extensive interactions with the CPU 120. second computational cycle; and 2. Minimizing data transfer between the expansion card 40 an accelerator controller, operably coupled to the accelera 180 and the system bus 200. All of the information needed to tor memory and the central processing unit, to transfer perform the simulations will be stored on the expansion card the at least the portion of the input data into the accel 180 and all simulations will take place on it. Furthermore, erator memory before the first computational cycle, to whatever data transfer remains necessary will take place in transfer the first output data from the accelerator parallel with the computation, thus reducing the impact of this 45 memory to the main memory during the second compu transfer on the performance. tational cycle, to direct the second output data into the 3. New way to execute GPU programs (shaders). Previ second partition during the second computational cycle, ously, the CPU 120 had full control over the order of shaders and to Swap the first pointer and the second pointer at the execution and was required to produce specific commands on conclusion of the second computational cycle Such that every cycle to tell the GPU 240 which shader to use. With the 50 the second output data becomes an input for a third invention disclosed herein, shaders will initially be stored on computational cycle of the plurality of computational the shader memory bank 210 on the expansion card 180 and cycles. will be sent to the GPU 240 for execution by the general 2. The computer system as claimed in claim 1, wherein the purpose controller 220 located on the expansion card. accelerator controller is configured to dictate an order of 4. Multiple parallelisms. The GPU 240 is inherently paral 55 execution of instructions to the at least one graphics process lel and is well suited to perform parallel computations. In ing unit. parallel with the GPU 240 performing the next calculation, 3. The computer system as claimed in claim 1, wherein the the controller 220 is uploading the data from the previous accelerator controller comprises an interface controller to calculation into main memory 130. Furthermore, the CPU communicate with the central processing unit over a bus of 120 at the same time uses uploaded previous results to save 60 the computer system. them onto disk 140 and to display them on the screen through 4. The computer system as claimed in claim 1, wherein the the system bus 200. accelerator memory comprises a texture memory bank to 5. Reuse of existing and affordable technology. All hard store the at least the portion of the input data and the first ware used in the invention and mentioned here-in are based on output data and a shader memory bank to store instructions currently available and reliable components. Further advance 65 for performing a set of operations to be performed on the at of these components will provide straightforward improve least the portion of the input data by the at least one graphic ments of the invention. processing unit. Case 7:26-mc-00318-LS Document 6-3 Filed 08/18/26 Page 15 of 15 US 8,648,867 B2 15 16 5. The computer system as claimed in claim 4, wherein the bank to store instructions for processing operations to be texture memory is partitioned into the first partition, the sec performed on the input data by the at least one graphic pro ond partition, a third partition to store internal variables, a cessing unit. fourth partition to store data textures used as input at a par 13. The accelerator system as claimed in claim 12, wherein ticular computation cycle of the plurality of computational 5 the texture memory is partitioned into a first partition to store cycles. the input data, a second partition to store internal variables, a 6. The computer system as claimed in claim 1, wherein the third partition to store data textures used as input at a particu accelerator controller inputs the at least the portion of the lar computation cycle of the numerical simulation, and a input data and a series of instructions into the at least one fourth partition to store the output data. graphic processing unit, wherein the at least one graphics 10 processing unit then executes the instructions on the at least the14. The accelerator system as claimed in claim 9, wherein accelerator controller is configured to perform successive the portion of the input data. computational cycles of the numerical simulation by feeding 7. The computer system as claimed in claim 1, wherein the the output data generated accelerator controller comprises a set of instructions stored ing unit from a previous by the at least one graphic process computational cycle and the input on a memory. 15 8. The computer system as claimed in claim 1, wherein the data for a next computational cycle into the at least one at least the portion of the input data represents an initial graphic processing unit. condition of the numerical simulation. 15. The accelerator system as claimed in claim 9, wherein 9. An accelerator system for a computer system performing the accelerator controller comprises a set of instructions a numerical simulation, the accelerator system comprising: 20 stored in a memory. 16. A method for performing a numerical simulation on at least one graphics processing unit to generate output data input data in a computer system including a central process by performing at least one computation during a first computational cycle of the numerical simulation; ing unit and an accelerator, the method comprising: an accelerator memory, operably coupled to the at least one receiving, by an accelerator, first input data from the central processing unit; graphics processing unit, to store data used to perform 25 transferring, the at least one computation; and by an accelerator controller, the first input an accelerator controller, operably coupled to the accelera data into a first partition, referenced by first pointer. ofan tor memory and the at least one graphics processing unit, accelerator memory before a first computational cycle of to execute: the numerical simulation; (i) a computational stream controlling performance of 30 performing, by at least one graphics processing unit during the first computational cycle, at least one calculation on the at least one computation by the at least one graph the first portion of the input data as to generate first ics processing unit; output data; (ii) an output stream controlling transfer of the output storing, by the accelerator controller, the first output data data from the at least one graphics processing unit to into a second partition, referenced by a second pointer, the accelerator memory during the first computational 35 of the accelerator memory; and cycle; and (iii) an input stream controlling transfer of input data to Swapping the first pointer with the second pointer at the end the accelerator memory for use by the at least one of the first computational cycle, such that the first output graphics processing unit during a second computa data becomes an input for a second computational cycle of the numerical simulation. tional cycle of the numerical simulation. 40 10. The accelerator system as claimed in claim 9, wherein 17. The method as claimed inclaim 16, further comprising: the accelerator controller is configured to dictate an order of sending, by the accelerator controller, instructions for per execution of instructions to the at least one graphics process forming the at least one calculation to the at least one ing unit. graphics processing unit. 11. The accelerator system as claimed in claim 9, wherein 45 18. The method as claimed in claim 16, further comprising: the accelerator controller is configured to send instructions to partitioning the accelerator memory into the first partition, the at least one graphics processing unit, wherein the at least the second partition, a third partition to store internal one graphics processing unit is configured to execute the Variables, and a fourth partition to store data used as instructions on the input data, and wherein the accelerator input at a particular computation cycle of the numerical simulation. controller is configured to transfer the output data from the 50 19. The method of claim 16, further comprising: accelerator memory to a main memory during the execution transferring, by the accelerator controller, the first output of the instructions by the at least one graphics processing unit. data to the main memory during the second computa 12. The accelerator system as claimed in claim 9, wherein tional cycle. the accelerator memory comprises a texture memory bank to Store the input data and the output data and a shader memory