jain.com
Public record. We host this document directly; the copy served here does not depend on any third party. Retrieved September 29, 2026.
Read the documentAlso at archive.org ↗

Neural AI, LLC v. Tesla Inc. — Entry #6: CORRECTED MOTION to Compel Compliance With Subpoena Served on Third Party Tesla, Inc

Case: Neural AI, LLC v. Tesla Inc. txwd · 7:26-cv-00318

filed August 17, 2026

What this document is

Docket entry #6 · filed August 18, 2026

CORRECTED MOTION to Compel Compliance With Subpoena Served on Third Party Tesla, Inc. by Neural AI, LLC. (Attachments: # 1 Affidavit Declaration of Tanner Laiche, # 2 Exhibit 1, # 3 Exhibit 2, # 4 Exhibit 3, # 5 Exhibit 4, # 6 Exhibit 5, # 7 Exhibit 6, # 8 Exhibit 7, # 9 Exhibit 8, # 10 Exhibit 9, # 11 Exhibit 10, # 12 Exhibit 11, # 13 Exhibit 12, # 14 Exhibit 13, # 15 Exhibit 14, # 16 Exhibit 15, # 17 Exhibit 16, # 18 Exhibit 17, # 19 Exhibit 18, # 20 Exhibit 19, # 21 Exhibit 20, # 22 Exhibit 21, # 23 Proposed Order)(Magni, Rocco) (Entered: 08/18/2026)

Who is involved

Why we have it

We follow this case because it names a company we track, although that company is not a party:

A free copy from the RECAP archive of federal court filings (mirrored at the Internet Archive), retrieved September 29, 2026. Federal court filings are public records.

URL
https://archive.org/download/gov.uscourts.txwd.1172927763/gov.uscourts.txwd.1172927763.6.3.pdf
Kind
court_filing
Publisher
RECAP
Retrieved
2026-09-29 05:58:24.933988-04:00
HTTP status
200
MIME
application/pdf
Bytes
1311511
SHA-256
73b0d285ba9c4ebd881284eb77b05d99bf3652c8b36eb239a4e1dbf57187be15

Document text

15 page(s), 78,530 characters, converted from the PDF's text layer · plain text.

Full text
Case 7:26-mc-00318-LS   Document 6-3   Filed 08/18/26   Page 1 of 15


                EXHIBIT

                              2


        Case 7:26-mc-00318-LS                               Document 6-3                  Filed 08/18/26                Page 2 of 15

                                                                                                       USOO8648867B2


(12) United States Patent                                                          (10) Patent No.:                  US 8,648,867 B2
       Gorchetchnikov et al.                                                       (45) Date of Patent:                        Feb. 11, 2014
(54)   GRAPHIC PROCESSORBASED                                                   (56)                    References Cited
       ACCELERATOR SYSTEMAND METHOD
                                                                                               U.S. PATENT DOCUMENTS
(75) Inventors: Anatoli Gorchetchnikov, Belmont, MA                                    5,388,206 A *    2/1995 Poulton et al. ................ 345,505
                (US); Heather Marie Ames, South                                  2005, 0166042 A1*      7, 2005 Evans ............           T13,150
                Boston, MA (US); Massimiliano                                    2007/0052713 A1* 3/2007 Chung et al.                     ... 34.5/5O1
                Versace, South Boston, MA (US);                                  2007/0279429 A1* 12/2007 Ganzer ......................... 345,582
                     Fabrizio Santini, Jamaica Plain, MA
               (US)                                                             * cited by examiner
(73) Assignee: Neurala LLC, Boston, MA (US)                                    Primary Examiner — Maurice L. McDowell, Jr.
(*) Notice: Subject to any disclaimer, the term of this                        (57)                      ABSTRACT
               patent is extended or adjusted under 35                         An accelerator system is implemented on an expansion card
               U.S.C. 154(b) by 1030 days.                                     comprising a printed circuit board having (a) one or more
                                                                               graphics processing units (GPU), (b) two or more associated
(21) Appl. No.: 11/860,254                                                     memory banks (logically or physically partitioned), (c) a
(22) Filed:          Sep. 24, 2007                                             specialized controller, and (d) a local bus providing signal
                                                                               coupling compatible with the PCI industry standards (this
(65)                     Prior Publication Data                                includes but is not limited to PCI-Express, PCI-X, USB 2.0,
                                                                               or functionally similar technologies). The controller handles
       US 2008/O 117220 A1              May 22, 2008                           most of the primitive operations needed to set up and control
            Related U.S. Application Data                                      GPU computation. As a result, the computer's central pro
                                                                               cessing unit (CPU) is freed from this function and is dedicated
(60) Provisional application No. 60/826,892, filed on Sep.                     to other tasks. In this case a few controls (simulation start and
       25, 2006.                                                               stop signals from the CPU and the simulation completion
                                                                               signal back to CPU), GPU programs and input/output data are
(51)   Int. C.                                                                 the information exchanged between CPU and the expansion
       G06F 5/00                    (2006.01)                                  card. Moreover, since on every time step of the simulation the
(52)   U.S. C.                                                                 results from the previous time step are used but not changed,
       USPC ........................................... 345/501; 34.5/503      the results are preferably transferred back to CPU in parallel
(58)   Field of Classification Search                                          with the computation.
       USPC .................................................. 345/501,503
       See application file for complete search history.                                       19 Claims, 5 Drawing Sheets

                      140                200
                                                                                                8O                                   250
                                 -4


                                                                             SHADERE.K.
                                                                                     MEMORY

                                                                                                                       50

                                                           O                                                             EXRE
                                                           -                   icon TROLLER w                           NEMORY

                                                           C
                                                           -


                                                                                                 2


   Case 7:26-mc-00318-LS     Document 6-3        Filed 08/18/26    Page 3 of 15


U.S. Patent        Feb. 11, 2014        Sheet 1 of 5              US 8,648,867 B2


              7.


                                          $3.


                                   s


                                       FG,


   Case 7:26-mc-00318-LS        Document 6-3        Filed 08/18/26    Page 4 of 15


U.S. Patent           Feb. 11, 2014        Sheet 2 of 5              US 8,648,867 B2


    g
                                      i.
              :
                  i
                  :                   r
              s
              al                 e3esiasti Od


    s

   3


   s


     Case 7:26-mc-00318-LS                                Document 6-3                  Filed 08/18/26                   Page 5 of 15


U.S. Patent                          Feb. 11, 2014                          Sheet 3 of 5                              US 8,648,867 B2


                                                    300
 CPU 120                              Start
                                                                                                 304
                                                                                                                Expansion Cardiso                     :
  iaia fict                               iss: Risitio:                                                          &::::::::8:::::::::
  Sires    301                                                                                                   Se::303
                                                                                                                                                320
            Disk:O                 Graphic User                                                                                 Controller
                                     Interface                                                                                 initialization
          initialization           Initialization


                                                                                                                                                325
                                      User                                        input Parser                                External input
                                                                                    Texture                               textures from RAM to
                                   interaction                                     Generator                              texture memory bank
                                                                                                                                                326   :
                                                                                Population Parser                          Population shader
                                                                                Shader Generato                           binaries from RAM to
                                                                                  and Compiler                            shader memory bank


                                                            Simulation
                                                           Initialization


                                                                                                                                                330


                                                            Progress                             Input Parser
                                                             monitor                               Texture
                                                                                                  Generator                  Computation
                                                                                                                              (see Fig. 4)
          Data output
            to Disk
                                                                                  Output Data            Data
                                                                                 ACCumulation                   issils
                                                                                     in RAM

              Last
           iteration
                           318 :

                                                             Yes
                                                                                                 350


                                                                     3


   Case 7:26-mc-00318-LS                  Document 6-3                    Filed 08/18/26                Page 6 of 15


U.S. Patent                   Feb. 11, 2014               Sheet 4 of 5                                US 8,648,867 B2


           Expansion Card 180

                                                    Simulation
                                                       start

                   iDesi: {34}{ut :   {iciplisiii) 2.                                                  iaia iiii:
                   Si:Sires:          Siii:Sire::                                                      Silisirai?
                   402                403                                                              404
                                                Input textures from
                                                  texture memory
                                                   bank to GPU


                                                  Shaders from
                                                 shader memory
                                                  bank to GPU


                                                                           ... ...........


                                                     Shader
                                                    execution


                                                 Output texture
                                                upload to texture
                                                 memory bank

                                                                                              New external input            Data
                                                                                             textures from RAM told -----------


         Wait for swap                             Swap input                                   Wait for swap
         of input/output                           and output                                  of input/output
        texture pointers                         texture pointers                              texture pointers

                                                                                                                            475
        Input textures from                                                                             NO           Last
         texture memory                                                                                            iteration?
          bank to RAM
                                                                                                                   Yes


            iteration?
            Yes                                                  490
                                                Wait for all three
                                                   streams of
                                               execution to finish

                                                                    499


                                              EIG 4


   Case 7:26-mc-00318-LS                Document 6-3    Filed 08/18/26    Page 7 of 15


U.S. Patent                   Feb. 11, 2014    Sheet 5 of 5              US 8,648,867 B2


                                         (p)


              ··*,?)&x<!- *


        Case 7:26-mc-00318-LS                      Document 6-3              Filed 08/18/26             Page 8 of 15


                                                     US 8,648,867 B2
                              1.                                                                2
         GRAPHIC PROCESSORBASED                                  embodiment, the context is initialized within a computational
      ACCELERATOR SYSTEMAND METHOD                               thread. This creates complications, however, in the interac
                                                                 tion between the user interface thread that changes param
                RELATED APPLICATIONS                             eters of simulations and the computational thread that uses
                                                                 these parameters.
   This application claims the benefit under 35 USC 119(e) of       A solution as proposed here is an implementation of the
U.S. Provisional Application No. 60/826,892, filed on Sep. computational stream of execution inhardware, so that thread
25, 2006, which is incorporated herein by reference in its and context initialization are replaced by hardware initializa
entirety.                                                        tion. This hardware implementation includes an expansion
                                                              10
          BACKGROUND OF THE INVENTION
                                                                 card comprising a printed circuitboard having (a) one or more
                                                                 graphics processing units, (b) two or more associated
   Graphics Processing Units (GPUs) are found in video memory    a
                                                                          banks that are logically or physically partitioned, (c)
                                                                   specialized controller, and (d) a local bus providing signal
adapters (graphic cards) of most personal computers (PCs), coupling compatible          with the PCI industry standards (this
Video game consoles, workstations, etc. and are considered includes but is not limited       to PCI-Express, PCI-X, USB 2.0,
highly parallel processors dedicated to fast computation of
graphical content. With the advances of the computer and         or functionally similar technologies).  The controller handles
console gaming industries, the need for efficient manipula most of the primitive operations needed to set up and control
tion and display of 3D graphics has accelerated the develop GPU computation. As a result, the CPU is freed from this
ment of GPUs.                                                       function and is dedicated to other tasks. In this case a few
   In addition, manufacturers of GPUs have included general controls (simulation start and stop signals from the CPU and
purpose programmability into the GPU architecture leading the simulation completion signal back to CPU), GPU pro
to the increased popularity of using GPUs for highly paral grams and input/output data are the information exchanged
lelizable and computationally expensive algorithms outside between CPU and the expansion card. Moreover, since on
of the computer graphics domain. When implemented on 25 every time step of the simulation the results from the previous
conventional video card architectures, these general purpose time step are used but not changed, the results are preferably
GPU (GPGPU) applications are not able to achieve optimal transferred back to CPU in parallel with the computation.
performance, however. There is overhead for graphics-related            In general, according to one aspect, the invention features
features and algorithms that are not necessary for these non a computer system. This system comprises a central process
Video applications.                                               30
                                                                     ing unit, main memory accessed by the central processing
             SUMMARY OF THE INVENTION                                unit, and a video system for driving a video monitor in
                                                                     response to the central processing unit as is common. The
   Numerical simulations, e.g., finite element analysis, of computer system further comprises an accelerator that uses
large systems of similar elements (e.g. neural networks, 35 input data from and provides output data to the central pro
genetic algorithms, particle systems, mechanical systems) are cessing unit. This accelerator comprises at least one graphics
one example of an application that can benefit from GPGPU processing unit, accelerator memory for the graphic process
computation. During numerical simulations, disk and user ing unit, and an accelerator controller that moves the input
input/output can be performed independently of computation data into the at least one graphics processing unit and the
because these two processes require interactions with periph 40 accelerator memory to generate the output data.
eral hardware (disk, Screen, keyboard, mouse, etc) and put              In the preferred, the central processing unit transfers the
relatively low load on the central processing unit/system input data for a simulation to the accelerator, after which the
(CPU). Complete independence is not desirable, however; accelerator executes simulation computations to generate the
user input might affect how the computation is performed and output data, which is transferred to the central processing
even interrupt it if necessary. Furthermore, the user output 45 unit. Preferably, the accelerator controller dictates an order of
and the disk output are dependent on the results of the com execution of instructions to the at least one graphics process
putation. A reasonable solution would be to separate input/ ing unit. The use of the separate controller enables data trans
output into threads, so that it is interacting with hardware fer during execution Such that the accelerator controller trans
occurs in parallel with the computation. In this case whatever fers output data from the accelerator memory to main
CPU processing is required for input/output should be 50 memory of the central processing unit.
designed so that it provides the synchronization with compu             In the preferred embodiment, the accelerator controller
tation.                                                              comprises an interface controller that enables the accelerator
   In the case of GPGPU, the computation itself is performed to communicate over a bus of the computer system with the
outside of the CPU, so the complete system comprises three central processing unit.
“peripheral components: user interactive hardware, disk 55 In general according to another aspect, the invention also
hardware, and computational hardware. The central process features an accelerator system for a computer system, which
ing unit (CPU) establishes communication and synchroniza comprises at least one graphics processing unit, accelerator
tion between peripherals. Each of the peripherals is prefer memory for the graphic processing unit and an accelerator
ably controlled by a dedicated thread that is executed in controller for moving data between the at least one graphics
parallel with minimal interactions and dependencies on the 60 processing unit and the accelerator memory.
other threads.                                                          In general according to another aspect, the invention also
   A GPU on a conventional video card is usually controlled features a method for performing numerical simulations in a
through OpenGL, DirectX, or similar graphic application computer system. This method comprises a central process
programming interfaces (APIs). Such APIs establish the con ing unit loading input data into an accelerator System from
text of graphic operations, within which all calls to the GPU 65 main memory of the central processing unit and an accelera
are made. This context only works when initialized within the tor controller transferring the input data to a graphics process
same thread of execution that uses it. As a result, in a preferred ing unit with instructions to be performed on the input data.


        Case 7:26-mc-00318-LS                     Document 6-3               Filed 08/18/26             Page 9 of 15


                                                    US 8,648,867 B2
                             3                                                                     4
The accelerator controller then transfers output data gener         PCI-X, or any other functionally similar technology (depend
ated by the graphic processing unit to the central processing       ing upon the availability on the motherboard 110). An exter
unit as output data.                                                nal version GPU accelerator is also a possible implementa
   The above and other features of the invention including          tion. In this example, the external GPU accelerator is
various novel details of construction and combinations of 5 connected to the motherboard 110 through USB-2.0, IEEE
parts, and other advantages, will now be more particularly          1394 (Firewire), or similar external/peripheral device inter
described with reference to the accompanying drawings and face.
pointed out in the claims. It will be understood that the par          The CPU 120 and the system memory 130 on the mother
ticular method and device embodying the invention are board 110 and the mass data storage system 140 are prefer
shown by way of illustration and not as a limitation of the 10 ably independent of the expansion card 180 and only com
invention. The principles and features of this invention may municate with each other and the expansion card 180 through
be employed in various and numerous embodiments without the system bus 200 located in the motherboard 110. A system
departing from the scope of the invention.                          bus 200 in current generations of computers have bandwidths
       BRIEF DESCRIPTION OF THE DRAWINGS                         15 from 3.2 GB/s (Pentium 4 with AGTL+, Athlon XP with
                                                                    EV6) to around 15 GB/s (Xeon Woodcrest with AGTL+,
   In the accompanying drawings, reference characters refer Athlon 64/Opteron with Hypertransport), while the local bus
to the same parts throughout the different views. The draw has maximal peak data transfer rates of 4GB/s (PCI Express
ings are not necessarily to scale; emphasis has instead been 16) or 2 GB/s (PCI-X 2.0). Thus the local bus 190 becomes a
placed upon illustrating the principles of the invention. Of the bottleneck in the information exchange between the system
drawings:                                                           bus 200 and the expansion card 180. The design of the expan
   FIG. 1 is a schematic diagram illustrating a computer sys sion card and methods proposed herein minimizes the data
tem including the GPU accelerator according to an embodi transfer through the local bus 190 to reduce the effect of this
ment of the present invention;                                      bottleneck.
   FIG. 2 is block diagram illustrating the architecture for the 25 The system memory 130 is referred to as the main random
GPU accelerator according to an embodiment of the present access memory (RAM) in the description herein. However,
invention;                                                          this is not intended to limit the system memory 130 to only
   FIG. 3 is a block/flow diagram illustrating an exemplary RAM technology. Other possible computer storage media
implementation of the top level control of the GPU accelera include, but are not limited to ROM, EEPROM, flash
tor system;                                                      30 memory, or any other memory technology.
   FIG. 4 is a flow diagram illustrating an exemplary imple            In the illustrated example, the GPU accelerator system is
mentation of the bottom level control of the GPU accelerator        implemented on an expansion card 180 on which the one or
system that is used to execute the target computation; and          more GPU's 240 are mounted. It should be noted that the
   FIG. 5 is an example population of nine computational GPU accelerator system GPU 240 is separate from and inde
elements arranged in a 3x3 square and a potential packing 35 pendent of any GPU on the standard video card 150 or other
scheme for texture pixels, according to an implementation of Video driving hardware such as integrated graphics systems.
the present invention.                                              Thus the computations performed on the expansion card 180
                                                                    do not interfere with graphics display (including but not lim
    DETAILED DESCRIPTION OF THE PREFERRED                           ited to manipulation and rendering of images).
                     EMBODIMENTS                                40     Various brand of GPU are relevant. Under current technol
                                                                     ogy, GPUs based on the GeForce series from NVIDIA Cor
  The Hardware                                                       poration or the Catalyst series from ATI/Advanced Micro
   FIG. 1 shows a computer system 100 that has been con Devices, Inc.
structed according to the principles of the present invention.      The output to a video monitor 170 is preferably through the
   In more detail, the computer system 100 in one example is 45 video card 150 and not the GPU accelerator system 180. The
a standard personal computer (PC). However, this only serves video card 150 is dedicated to the transfer of graphical infor
as an example environment as computing environment 100 mation and connects to the motherboard 110 through a local
does not necessarily depend on or require any combination of bus 160 that is sometimes physically separate from the local
the components that are illustrated and described herein. In bus 190 that connects the expansion card 180 to the mother
fact, there are many other Suitable computing environments 50 board 110.
for this invention, including, but not limited to, workstations,    FIG. 2 is a block diagram illustrating the general architec
server computers, Supercomputers, notebook computers, ture of the GPU accelerator system and specifically the
hand-held electronic devices such as cell phones, mp3 play expansion card 180 in which at least one GPU 240 and asso
ers, or personal digital assistants (PDAs), multiprocessor sys ciated memories 210 and 250 are mounted. Electrical (signal)
tems, programmable consumer electronics, networks of any 55 and mechanical coupling with a local bus 190 provides signal
of the above-mentioned computing devices, and distributed coupling compatible with the PCI industry standards (this
computing environments that including any of the above includes but is not limited to PCI, PCI-X, PCI Express, or
mentioned computing devices.                                     functionally similar technology).
   In one implementation the GPU accelerator is imple               The GPU accelerator further preferably comprises one spe
mented as an expansion card 180 includes connections with 60 cifically designed accelerator controller 220. Depending
the motherboard 110, on which the one or more CPU's 120          upon the implementation, the accelerator controller 220 is
are installed along with main, or system memory 130 and field programmable gate array (FPGA) logic, or custom built
mass/non Volatile data storage 140. Such as hard drive or application-specific (ASIC) chip mounted in the expansion
redundant array of independent drives (RAID) array, for the card 180, and in mechanical and signal coupling with the
computer system 100. In the current example, the expansion 65 GPU 240 and the associated memories 210 and 250. During
card 180 communicates to the motherboard 110 via a local         initial design, a controller can be partially or even fully imple
bus 190. This local bus 190 could be PCI, PCI Express, mented in Software, in one example.


       Case 7:26-mc-00318-LS                    Document 6-3              Filed 08/18/26            Page 10 of 15


                                                   US 8,648,867 B2
                              5                                                               6
   The controller 220 commands the storage and retrieval of represented as a texture. The important difference between
arrays of data (on a conventional video card the arrays of data output variables and internal variables is their access.
are represented as textures, hence the term texture in this       Output variables are usually accessed by any element in the
document refers to a data array unless specified otherwise and system during every time step. The value of the output vari
each element of the texture is a pixel of color information), able that is accessed by other elements of the system corre
execution of GPU programs (on a conventional video card sponds to the value computed on the previous, not the current,
these programs are called shaders, hence the term shader in time step. This is realized by dedicating two textures to output
this document refers to a GPU program unless specified oth variables—one holds the value computed during the previous
erwise), and data transfer between the system bus 200 and the time step and is accessible to all computational elements
expansion card 180 through the local bus 190 which allows 10 during the current time step, another is not accessible to other
communication between the main CPU 120, RAM 130, and             elements and is used to accumulate new values for the vari
disk 140.                                                          able computed during the current time step. In-between time
   Two memory banks 210 and 250 are mounted on the expan steps these two textures are Switched, so that newly accumu
sion card 180. In some example, these memory banks sepa lated values serve as accessible input during the next time
rated in the hardware, as shown, or alternatively implemented 15 step, while the old input is replaced with new values of the
as a single, logically partitioned memory component.               variable. This Switch is implemented by Swapping the address
   The reason to separate the memory into two partitions 210 pointers to respective textures as described in the System and
250 stems from the nature of the computations to which the Framework section.
GPU accelerator system is applied. The elements of compu              Internal variables are computed and used within the same
tation (computational elements) are characterized by a single computational element. There is no chance of a race condition
output variable. Such computational elements often include in which the value is used before it is computed or after it has
one or more equations. Computational elements are same or already changed on the next time step because within an
similar within a large population and are computed in paral element the processing is sequential. Therefore, it is possible
lel. An example of Such a population is a layer of neurons in to render the new value of internal variable into the same
an artificial neural network (ANN), where all neurons are 25 texture where the old was read from in the texture memory
described by the same equation. As a result, some data and bank. Rendering to more than one texture from a single
most of the algorithms are common to all computational shader is not implemented in current GPU architectures, so
elements within population, while most of the data and some computational elements that track internal variables would
algorithms are specific for each equation. Thus, one memory, have to have one shader per variable. These shaders can be
the shader memory bank 210, is used to store the shaders 30 executed in order with internal variables computed first, fol
needed for the execution of the required computations and the lowed by output variables.
parameters that are common for all computational elements             Further savings of texture memory is achieved through
and is coupled with the controller 220 only. The second using multiple color components per pixel (texture element)
memory, the texture memory bank 250, is used to store all the to hold data. Textures can have up to four color components
necessary data that are specific for every computational ele 35 that are all processed in parallel on a GPU. Thus, to maximize
ment (including, but not limited to, input data, output data, the use of GPU architecture it is desirable to pack the data in
intermediate results, and parameters) and is coupled with Sucha way that all four components are used by the algorithm.
both the controller 220 and the GPU 240.                           Even though each computational element can have multiple
   The texture memory bank 250 is preferably further parti variables, designating one texture pixel per element is inef
tioned into four sections. The first partition 250a is designed 40 fective because internal variables require one texture and
to hold the external input data patterns. The second partition output variables require two textures. Furthermore, different
250b is designed to hold the data textures representing inter element types have different numbers of variables and unless
nal variables. The third partition 250c is designed to hold the this number is precisely a multiple of four, texture memory
data textures used as input at a particular computation step on can be wasted.
the GPU 240. The fourth partition 250d holds the data tex 45 A more reasonable packing scheme would be to pack four
tures used to accommodate the output of a particular compu computational elements into a pixel and have separate tex
tational step on the GPU240. This partitioning scheme can be tures for every variable associated with each computational
done logically, does not require hardware implementation. element. In this case the packing scheme is identical for all
Also the partitioning scheme is also altered based on new textures, and therefore can be accessed using the same algo
designs or needs of the algorithms being employed. The rea 50 rithm. Several ways to approach this packing scheme are
son for this partitioning is further explained in the Data Orga outlined here. An example population of nine computational
nization section, below.                                           elements arranged in a 3x3 square (FIG.5a) can be packed by
   A local bus interface 230 on the controller 220 serves as a     element (FIG.5b), by row (FIG.5c), or by square (FIG. 5d).
driver that allows the controller 220 to communicate through          Packing by element (FIG.5b) means that elements 1.2.3.4
the local bus 190 with the system bus 200 and thus the CPU 55 go into first pixel; 5.6.7.8 go into second pixel; 9 goes into
120 and RAM 130. This local bus interface 230 is not               third pixel. This is the most compact scheme, but not conve
intended to be limited to PCI related technology. Other driv nient because the geometrical relationship is not preserved
ers can be used to interface with comparable technology as a during packing and its extraction depends on the size of the
local bus 190.                                                     population.
   Data Organization                                            60    Packing by row (column; FIG. 5c) means that elements
   Each computational element discussed above has output 1.2.3 go into pixel (1,1); 3.45 go into pixel (2,1), 7.8.9 go into
variables that affect the rest of the system. For example in the pixel (3,1). With this scheme the element’sy coordinate in the
case of a neural network it is the output of a neuron. A population is the pixel’s y coordinate, while the elements x
computational element also usually has several internal vari coordinate in the population is the pixel’s X coordinate times
ables that are used to compute output variables, but are not 65 four plus the index of color component. Five by five popula
exposed to the rest of the system, not even to other elements tions in this case will use 2x5 texture, or 10 pixels. Five of
of the same population, typically. Each of these variables is these pixels will only use one out of four components, so it


       Case 7:26-mc-00318-LS                     Document 6-3                Filed 08/18/26            Page 11 of 15


                                                    US 8,648,867 B2
                                7                                                                  8
wastes 37.5% of this texture. 25x1 population will use 6x1         external inputs. It specifies which equations should have their
texture (six pixels) and will waste 12.5% of it.                   output saved to disk and/or displayed on the screen. It allows
   Packing by square (FIG. 5d) means that elements 1,2,4,5 the user to start and stop the simulation. And it performs
go into pixel (1,1); 3.6 go into pixel (1,2); 7.8 go into pixel standard interface functions such as file loading and saving,
(2,1), and 9 goes into pixel (2.2). Both the row and the column interactive help, general preferences and others.
of the element are determined from the row (column) of the            The user interaction 305 directs the CPU 120 to acquire the
pixel times two plus the second (first) bit of the color com new external input textures needed (this includes but is not
ponent index. Five by five populations in this case will use limited to loading from disk 140 or receiving them in real time
3x3 texture, or 9 pixels. Four of these pixels will only use two from a recording device), parses them if necessary 309, and
out of four components, and one will only use one compo 10 initializes their transfer to the expansion card 180, where they
nent, so it wastes 34.4% of this texture. This is more advan       are stored 325 in the texture memory bank 250 by the con
tageous than packing by row, since the texture is Smaller and troller 220. The user interaction 305 also directs the CPU 120
the waste is also lower. 25x1 population on the other hand will to parse populations of elements that will be used in the
use 13x1 texture (thirteen pixels) and waste D-50% of it, which simulation, convert them to GPU programs (shaders), com
is much worse than packing by row.                              15 pile them 310, and initializes their transfer to the expansion
   In order to eliminate waste altogether the population card 180, where they are stored 326 in the shader memory
should have even dimensions in the square packing, and it bank 210 by the controller 220. This operation is accompa
should have a number of columns divisible by four in row nied by the upload 309 of the initial data into the input parti
packing. Theoretically, the chances are approximately tion of the texture memory bank 250, and stores the shader
equivalent for both of these cases to occur, so the particular order of execution in the controller 220. The user can perform
task and data sizes should determine which packing scheme is operations 309 and 310 as many times as necessary prior to
preferable in each individual case.                                starting the simulation or between simulations.
   The System and Framework                                           The editing of the system between simulations is difficult
   FIG.3 shows an exemplary implementation of the top level to accomplish without the hardware implementation of the
system and method that is used to control the computation. It 25 computational thread suggested herein. The system of equa
is a representation of one of several ways in which a system tions (computational elements) is represented by textures that
and method for processing numerical techniques can be track variables plus shaders that define processing algo
implemented in the invention described herein and so the rithms. As mentioned above, textures, shaders and other
implementation is not intended to be limited to the following graphics related constructs can only be initialized within the
description and accompanying figure.                            30 rendering context, which is thread specific. Therefore tex
   The method presented herein includes two execution tures and shaders can only be initialized in the computational
streams that run on the CPU 120 User Interaction Stream            thread.
302 and Data Output Stream 301. These two streams prefer             Network editing is a user-interactive process, which
ably do not interact directly, but depend on the same data according to the scheme suggested above happens in the User
accumulated during simulations. They can be implemented as 35 Interaction Stream 302. The simulation software thus has to
separate threads with shared memory access and executed on take the new parameters from the User Interaction Stream
different CPUs in the case of multi-CPU computing environ 302, communicate them to the Computational Stream 303
ment. The third execution stream—Computational Stream and regenerate the necessary shaders and textures. This is
303 runs on the GPU accelerator of the expansion card 180 hard to accomplish without a hardware implementation of the
and interacts with the User Interaction Stream 302 through 40 Computational Stream 303. The Computational Stream 303
initialization routines and data exchange in between simula is forked from the User Interaction Stream and it can access
tions. The Computational Stream 303 interacts with the User the memory of the parent thread, but the reverse communica
Interaction Stream and the Data Output Stream through syn tion is harder to achieve. The controller 220 allows operations
chronization procedures during simulations.                        309 and 310 to be performed as many times as necessary by
   The crucial feature of the interaction between the User 45 providing the necessary communication to the User Interac
Interaction Stream 302 and the Computational Stream 303 is tion Stream 302.
the shift of priorities. Outside of the simulation, the system       After execution of the input parser texture generation 309
100 is driven by the user input, thus the User Interaction and population parser shader generator and compiler 310 are
Stream 302 has the priority and controls the data exchange performed at least once, the user has the option to initialize the
304 between streams. After the user starts the simulation, the 50 simulation 311. During this initialization the main control of
Computational Stream 303 takes the priority and controls the the framework is transferred to the GPU accelerator systems
data exchange between streams until the simulation is fin accelerator controller 220 and computation 330 is started (see
ished or interrupted 350.                                          FIG. 4; 420). The user retains the ability to interrupt the
   The user starts 300 the framework through the means of an simulation, change the input, or to change the display prop
operating system and interacts with the Software through the 55 erties of the framework, but these interactions are queued to
user interaction section 305 of the graphic user interface 306 be performed at times determined by the controller-driven
executed on the CPU 120. The start 300 of the implementa data exchange 314 and 316 to avoid the corruption of the data.
tion begins with a user action that causes a GUI initialization      The progress monitor 312 is not necessary for perfor
307, Disk input/output initialization 308 on the CPU 120, and mance, but adds convenience. It displays the percentage of
controller initialization 320 of the GPU accelerator on the 60 completed time steps of the simulation and allows the user to
expansion card 180. GUI initialization includes opening of plan the schedule using the estimates of the simulation wall
the main application window and setting the interface tools clock times. Controller-driven data exchange 314 updates the
that allow the user to control the framework. Disk I/O initial     display of the results 313. Online screen output for the user
ization can be performed at the start of the framework, or at selected population allows the user to monitor the activity and
the start of each individual simulation.                        65 evaluate the qualitative behavior of the network. Simulations
   The user interaction 305 controls the setting and editing of with unsatisfactory behavior can be terminated early to
the computational elements, parameters, and sources of change parameters and restart. Controller-driven data


        Case 7:26-mc-00318-LS                         Document 6-3                Filed 08/18/26                Page 12 of 15


                                                         US 8,648,867 B2
                                                                                                            10
exchange 314 also drives the output of the results to disk317.      element-specific.        Element      independent     objects include Sub
Data output to disk for convenience can be done on an ele components of TEquation and objects that describe how to
ment per file basis. A suggested file format includes a leftmost handle interdependencies between variables implemented
column that displays a simulated time for each of the simu through derivatives of TGate class.
lation steps and Subsequent columns that display variable              Element-specific data is held in TElement objects. These
values during this time step in all elements with identical objects hold references to TEquation and a set of TGate
equations (e.g. all neurons in a layer of a neural network).        objects. There is one TElement perpopulation, but the size of
   Controller-driven data exchange or input parser texture data arrays within this object corresponds to population size.
generator 316 allows the user to change input that is generated All TElement objects have to be added to the TSimulator list
on the fly during the simulation. This allows the framework 10 of elements by calling TSimulator::addUnit() method from
monitoring of the input that is coming from a recording TPopulation::fillElements().
device (video camera, microphone, cell recording electrode,            Finally, TPopulation::fillElements( ) should contain a set
etc) in real time. Similar to the initial input parser 309, it of TElement:add Dependency() calls for each element.
preprocesses the input into a universal format of the data array Each of these calls sets a corresponding dependency for every
Suitable for texture generation and generates textures. Unlike 15 TGate object. Here TGate object holds element independent
the initial parser 309, here the textures are transferred to part of dependency and TElement::add Dependency() sets
hardware not whenever ready but upon the request of the element-specific details.
controller 220.                                                        System provided TPopulation handles the output of com
   The controller 220 also drives the conditional testing 315 putational elements, both when they need to exchange the
and 318 informs the CPU-bound streams whether the simu              data and when they need to output it to disk. User implemen
lation is finished. If so, the control returns to the User Inter    tation of TPopulation derivative can add screen output.
action Stream. The user then can change parameters or inputs           Listing 1 is an example code of the user program that uses
(309 and 310), restart the simulation (311) or quit the frame a recurrent competitive field (RCF) equation:
work (390).
   SANNDRA (Synchronous Artificial Neuronal Network 25
Distributed Runtime Algorithm: http://www.kinness.net/
Docs/SANNDRA/html) was developed to accelerate and uint16floatt wm= 3,compet          h = 3;
optimize processing of numerical integration of large non static                             = 0.5;
homogenous systems of differential equations. This library is static       float m persist = 1.0;
                                                                    class TCablePopRCF : public TPopulation
fully reworked in its version 2.X.X to support multiple com 30
putational backends including those based on multicore TEq RCF*gate1:              m equation;
CPUs, GPUs and other processing systems. GPU based back TGate*m     TGate* m gate2;
end for SANNDRA-2.x.x can serve as an example practical void createCatingStructure()
software implementation of the method and architecture
described above and pictorially represented in FIG. 3.           35 m gate1 = new TGate(O);
   To use SANNDRA, the application should create a TSimu m gate2 = new TGate(1):
lator object either directly or through inheritance. This object void createUnitStructure(TBasicUnitu)
will handle global simulation properties and control the User {
                                                                      u->addO2OInputDependency(m gate1, O., O., 0.004, O., 0, 0);
Interaction Stream, Data Output Stream, and Computational u->addFullDependency(m                    gate2, population());
Stream. Through TSimulator:timestep(), TSimulator:out 40
fileInterval ( ), and TSimulator:outmode(), the application public: TCablePopRCF(): TPopulation(“compCPU RCF, w, h, true)
can set the time step of the simulation, the time step of disk { };
output, and the mode of the disk output. The external input ~TCablePopRCF()               {if(m equation) delete m equation;
                                                                         if(m gate1) delete m gate1;
pattern should be packed into a TPattern object and bound to
the simulation object through TSimulator:resetInputs( ) 45 bool if(m             gate2) delete m gate2:};
                                                                          fillElements(TSimulator sim);
method. TSimulator::simLength( ) sets the length of the }:
simulation.                                                              bool TCablePopRCF::fillElements(TSimulatior sim)
   The second step is to create at least one population of { equation = new TEq RCF (this, m compet, m persist);
equations (TPopulation object). Population holds one equa mcreateCatingStructure();
tion object TEquation. This object contains only a formula 50 for(size ti = 0; i < x.Size(); ++i)
and does not hold element-specific data, so all elements of the   for(size tj = 0; j < ySize(); ++)
population can share single TEquation.                            {
   The TEquation object is converted to a GPU program TElement             u = new TCPUElement(this, m equation, i,j):
                                                                sim->addUnit(u);
before execution. GPU programs have to be executed within createUnitStructure(u);
a graphical context, which is stream specific. TSimulator 55 return true:
creates this context within a Computational Stream, therefore
all programs and data arrays that are necessary for computa int
tion have to be initialized within Computational Stream. Con main()
structor of TPopulation is called from User Interaction //{ Input pattern generation (309 in FIG. 3)
Stream, so no GPU-related objects can be initialized in this 60 uint32 t pat = new uint32 twh;
COnStructOr.                                                     TRandom-float randGen (O);
   TPopulation::fillElements( ) is a virtual method designed     for(uint32 ti = 0; i < wh; ++i)
to overcome this difficulty. It is called from within the Com    pati = randGen.random ();
putational Stream after TSimulator::networkCreate( ) is          TPattern p = new TPattern (pat, w, h);
                                                                 if Setting up the simulation
called in the User Interaction Stream. A user has to override 65 TSimulator cableSim = new TSimulator(“data'); fi(308 and 320 in
TPopulation::fillElements( ) to create TEquation and other FIG. 3)
computation related objects both element independent and


         Case 7:26-mc-00318-LS                             Document 6-3          Filed 08/18/26               Page 13 of 15


                                                            US 8,648,867 B2
                                   11                                                               12
                               -continued                               need to perform the computations and initiates the upload 435
                                                                        of them onto the GPU 240. The GPU 240 can communicate
cableSim->timestep(0.05); fi(320 in FIG. 3)
cableSim->resetInputs(p); (325 in FIG. 3)                               directly with the texture memory bank 250 to upload the
cableSim->OutfileInterval(0.1); (308 in FIG. 3)                         appropriate texture to perform the computations. The control
cableSim->Outmode(SANNDRA::timefunc); (308 in FIG. 3)
cableSim->simLength(60.0); (320 in FIG. 3)
                                                                        ler 220 also pulls the first shader (known by the stored order)
if Preparing the population                                             from the shader memory bank 210 and uploads 450 it onto the
TPopulation* cablePop = new TCablePopRCF(); //(310 in FIG. 3)           GPU 240.
cableSim->networkCreate(); //(326 in FIG. 3)                              The GPU 240 executes the following operations in this
uint16 t user = 1;                                                     order: performs the computation (execution of the shader)
while(user)                                                         10
                                                                       470; tells the controller 220 that it is done with the computa
if(cableSim->simulationStart(true, 1)) (311 in FIG. 3)                  tions for the current shader; and after all shaders for this
exit(1):                                                               particular equation are executed sends 480 the output textures
std::cout-“Repeat?\n: //(305 in FIG. 3)
std::cin>user; //(305 in FIG. 3)                                       to the output portion of the texture memory bank 250. This
if(user == 1)                                                       15 cycle continues through all of the equations based on the
cableSim->networkReset(); //(305 in FIG. 3)                            branching step 482.
if cableSim)                                                              An example shader that performs fourth order Runge
delete cableSim; Also deletes cablePop and its internals               Kutta numerical integration is shown in Listing 2 using GLSL
exit(0);                                                                notation;
Listing 1.

   FIG. 4 is a detailed flow diagram illustrating a part of an            uniform sampler2DRect Variable;
                                                                          uniform float integration step;
exemplary implementation of the bottom level system and                   float halfstep = integration step*0.5;
method performed during the computation on the GPU accel 25               float fl. 6 step = integration stepf 6.0:
erator of the expansion card 180 and is a more detailed view              vec4 output = texture2DRect(Variable, gl TexCoord O.st);
                                                                          if define equation() here
of the computational box 330 in FIG. 3. FIG. 4 is a represen              vec4 rungekutta4(vec4 x)
tation of one of several ways in which a system and method
for processing numerical techniques can be implemented.                     const vecA k1 = equation(x);
                                                                            const vec4 k2 = equation(x + halfstep*k1);
   With systems of equations that have complex interdepen 30                const vec4 k3 = equation(x + halfstep*k2);
dencies it is likely that the variable in Some equation from a              const vecA k4 = equation(x + integration Step*k3);
previous time step has to be used by some other equation after              return fl. 6step*(k1 + 2.0* (k2 + k3) + k4);
the new values of this variable are already computed for new
time step. To avoid data confusion, the new values of variables           void main (void)
                                                                          {
should be rendered in a separate texture. After the time step is 35         output += rungekutta4(output);
completed for all equations, these new values should be cop                 gl FragColor = output;
ied over old values so that they are used as input during the             Listing 2.
next time step. Copying textures is an expensive operation,
computationally, but since the textures are referred to by
texture IDs (pointers), Swapping these pointers for input and 40 The shader in Listing 2 can be executed on conventional
output textures after each time step achieves the same resultat video card. Using the controller 220 this code can be further
a much lesser cost.                                                 optimized, however. Since the integration step does not
   In the hardware solution suggested herein, ID Swapping is change during the simulation, the step itself as well as the
equivalent to Swapping the base memory address for two halfstep and /6 of the step can be computed once per simula
partitions of the texture memory bank 250. They are swapped 45 tion, and updated in all shaders by a shader update procedures
485 during synchronization (485, 430, and 455) so that data
transfer 445 and the computation 435-487 proceeds immedi 310,326               discussed above.
ately and in parallel with data transfer as shown in FIG. 4. A computed theofmain
                                                                       After  all      the equations in the computational cycle are
hardware solution allows this parallelism through access of 220 can switch 485execution        the
                                                                                                         substream 403 on the controller
                                                                                                    reference     pointers of the input and
the controller 220 to the onboard texture memory bank 250. 50 output portions of the texture memory                   bank 250.
   The main computation and data exchange are executed by              The two other substreams of execution on the controller
the controller 220. It runs three parallel substreams of execu
tion: Computational Substream 403, Data Output Substream 220 are waiting (blocks 430 and 455, respectively) for this
402, and Data Input Substream 404. These streams are syn switch to begin their execution. The Data Input Substream
chronized with each other during the swap of pointers 485 to 55 404 is controlling 440 the input of additional data from the
the input and output texture memory partitions of the texture CPU 120. This is necessary in cases where the simulation is
memory bank 250 and the check for the last iteration 487. monitoring the changing input, for example input from a
Algorithmically, these two operations are a single atomic video camera or other recording device in the real time. This
operation, but the block diagram shows them as two separate substream uploads new external input from the CPU 120 to
blocks for clarity.                                              60 the texture memory bank 250 so it can be used by the main
   The Computational Substream 403 performs a computa computational Substream 403 on the next computational step
tional cycle including a sequential execution of all shaders and waits for the next iteration 475. The Data Output Sub
that were stored in the shader memory bank 210 using the stream 445 controls the output of simulation results to the
appropriate input and output textures. To begin the simulation CPU 120 if requested by the user. This substream uploads the
the controller 220 initializes three execution substreams 403, 65 results of the previous step to the main RAM 130 so that the
402, and 404. On every simulation step, the Computational CPU 120 can save them on disk 140 or show them on the
Substream 403 determines which textures the GPU 240 will            results display 313 and waits for the next iteration 460.


       Case 7:26-mc-00318-LS                         Document 6-3                Filed 08/18/26               Page 14 of 15


                                                        US 8,648,867 B2
                                  13                                                                     14
   Since the Computational Substream 403 determines the                    While this invention has been particularly shown and
timing of input 440 and output 445 data transfers, these data described with references to preferred embodiments thereof,
transfers are driven by the controller 220. To further reduce it will be understood by those skilled in the art that various
the data transfer overhead (and disk 140 overhead also) the changes in form and details may be made therein without
controller 220 initiates transfer only after selected computa departing from the scope of the invention encompassed by the
tional steps. For example, if the experimental data that is appended claims.
simulated was recorded every 10 milliseconds (msec) and the
simulation for better precision was computed every 1 mSec,                 What is claimed is:
then only every tenth result has to be transferred to match the            1. A computer system for performing a numerical simula
experimental frequency.                                              10 tion over a plurality of computational cycles including at least
   This solution stores two copies of output data, one in the a first computational cycle and a second computational cycle,
expansion card texture memory bank 250 and another in the the computer system comprising:
system RAM 130. The copy in the system RAM 130 is                          a central processing unit;
accessed twice: for disk I/O and screen visualization 313. An              a main memory, operably coupled to the central processing
alternative solution would be to provide CPU 120 with a 15                    unit, to store input data to be accessed by the central
direct read access to the onboard texture memory bank 250 by                  processing unit in performing the numerical simulation;
mapping the memory of the hardware onto a global memory                    a video system, operably coupled to the central processing
space. The alternative solution will double the communica                     unit, to drive a video monitor to display an indication of
tion through the local bus 190. Since the goal discussed herein               the numerical simulation in response to the computer
is reducing the information transfer through the local bus 190,               system performing the numerical simulation;
the former solution is favored.                                            an accelerator, operably coupled to the central processing
   The main stream substream 403 determines if this is the last               unit, to receive at least a portion of the input data from
iteration 487. If it is the last iteration, the controller 220 waits          the central processing unit and to provide first output
for the all of the execution substreams to finish 490 and then                data generated during the first computational cycle to the
returns the control to the CPU 120, otherwise it begins the 25                central processing unit after a conclusion of the first
next computational cycle.                                                     computational cycle, the accelerator comprising:
   This repeats through all of the computational cycles of the                at least one graphics processing unit to generate second
simulation.                                                                      output data, during the second computational cycle,
   Conclusion                                                                    by performing at least one calculation on the first
   This GPU accelerator system offers the following potential 30                 output data; and
advantages:                                                                   an accelerator memory, operably coupled to the at least
    1. Limited computations on the CPU 120. The CPU 120 is                       one graphic processing unit, the accelerator memory
only used for user input, sending information to the controller                  comprising:
220, receiving output after each computational cycle (or less                    a first partition, referenced by a first pointer, to store
frequently as defined by the user), writing this output to disk 35                  the first output data during the second computa
140, and displaying this output on the monitor 170. This frees                      tional cycle; and
the CPU 120 to execute other applications and allows the                         a second partition, referenced by a second pointer, to
expansion card to run at its full capacity without being slowed                     store the second output data generated during the
down by extensive interactions with the CPU 120.                                    second computational cycle; and
   2. Minimizing data transfer between the expansion card 40 an accelerator controller, operably coupled to the accelera
180 and the system bus 200. All of the information needed to                  tor memory and the central processing unit, to transfer
perform the simulations will be stored on the expansion card                  the at least the portion of the input data into the accel
180 and all simulations will take place on it. Furthermore,                   erator memory before the first computational cycle, to
whatever data transfer remains necessary will take place in                   transfer the first output data from the accelerator
parallel with the computation, thus reducing the impact of this 45            memory to the main memory during the second compu
transfer on the performance.                                                  tational cycle, to direct the second output data into the
   3. New way to execute GPU programs (shaders). Previ                        second partition during the second computational cycle,
ously, the CPU 120 had full control over the order of shaders                 and to Swap the first pointer and the second pointer at the
execution and was required to produce specific commands on                    conclusion of the second computational cycle Such that
every cycle to tell the GPU 240 which shader to use. With the 50              the second output data becomes an input for a third
invention disclosed herein, shaders will initially be stored on               computational cycle of the plurality of computational
the shader memory bank 210 on the expansion card 180 and                      cycles.
will be sent to the GPU 240 for execution by the general                   2. The computer system as claimed in claim 1, wherein the
purpose controller 220 located on the expansion card.                   accelerator controller is configured to dictate an order of
   4. Multiple parallelisms. The GPU 240 is inherently paral 55 execution of instructions to the at least one graphics process
lel and is well suited to perform parallel computations. In ing unit.
parallel with the GPU 240 performing the next calculation,                 3. The computer system as claimed in claim 1, wherein the
the controller 220 is uploading the data from the previous accelerator controller comprises an interface controller to
calculation into main memory 130. Furthermore, the CPU communicate with the central processing unit over a bus of
120 at the same time uses uploaded previous results to save 60 the computer system.
them onto disk 140 and to display them on the screen through               4. The computer system as claimed in claim 1, wherein the
the system bus 200.                                                     accelerator memory comprises a texture memory bank to
   5. Reuse of existing and affordable technology. All hard store the at least the portion of the input data and the first
ware used in the invention and mentioned here-in are based on           output data and a shader memory bank to store instructions
currently available and reliable components. Further advance 65 for performing a set of operations to be performed on the at
of these components will provide straightforward improve least the portion of the input data by the at least one graphic
ments of the invention.                                                 processing unit.


       Case 7:26-mc-00318-LS                        Document 6-3                Filed 08/18/26              Page 15 of 15


                                                       US 8,648,867 B2
                               15                                                                      16
   5. The computer system as claimed in claim 4, wherein the bank to store instructions for processing operations to be
texture memory is partitioned into the first partition, the sec performed on the input data by the at least one graphic pro
ond partition, a third partition to store internal variables, a cessing unit.
fourth partition to store data textures used as input at a par         13. The accelerator system as claimed in claim 12, wherein
ticular computation cycle of the plurality of computational 5 the texture memory is partitioned into a first partition to store
cycles.                                                              the input data, a second partition to store internal variables, a
   6. The computer system as claimed in claim 1, wherein the third partition to store data textures used as input at a particu
accelerator controller inputs the at least the portion of the lar computation cycle of the numerical simulation, and a
input data and a series of instructions into the at least one fourth partition to store the output data.
graphic processing unit, wherein the at least one graphics 10
processing unit then executes the instructions on the at least the14.       The accelerator system as claimed in claim 9, wherein
                                                                         accelerator   controller is configured to perform successive
the portion of the input data.                                       computational     cycles  of the numerical simulation by feeding
   7. The computer system as claimed in claim 1, wherein the the output data generated
accelerator controller comprises a set of instructions stored ing unit from a previous by              the at least one graphic process
                                                                                                    computational cycle and the input
on a memory.                                                      15
   8. The computer system as claimed in claim 1, wherein the data for a next computational cycle into the at least one
at least the portion of the input data represents an initial graphic processing unit.
condition of the numerical simulation.                                 15. The accelerator system as claimed in claim 9, wherein
   9. An accelerator system for a computer system performing the          accelerator controller comprises a set of instructions
a numerical simulation, the accelerator system comprising: 20 stored         in a memory.
                                                                        16. A method for performing a numerical simulation on
   at least one graphics processing unit to generate output data input      data in a computer system including a central process
      by performing at least one computation during a first
      computational cycle of the numerical simulation;               ing unit and an accelerator, the method comprising:
   an accelerator memory, operably coupled to the at least one         receiving, by an accelerator, first input data from the central
                                                                           processing unit;
      graphics processing unit, to store data used to perform 25 transferring,
      the at least one computation; and                                                 by an accelerator controller, the first input
   an accelerator controller, operably coupled to the accelera             data into a first partition, referenced by first pointer. ofan
      tor memory and the at least one graphics processing unit,            accelerator memory before a first computational cycle of
      to execute:                                                          the numerical simulation;
      (i) a computational stream controlling performance of 30 performing,             by at least one graphics processing unit during
                                                                           the first computational cycle, at least one calculation on
         the at least one computation by the at least one graph            the first portion of the input data as to generate first
         ics processing unit;                                              output data;
      (ii) an output stream controlling transfer of the output         storing,   by the accelerator controller, the first output data
         data from the at least one graphics processing unit to            into a second partition, referenced by a second pointer,
         the accelerator memory during the first computational 35          of the accelerator memory; and
         cycle; and
      (iii) an input stream controlling transfer of input data to      Swapping the first pointer with the second pointer at the end
         the accelerator memory for use by the at least one                of the first computational cycle, such that the first output
         graphics processing unit during a second computa                  data becomes an input for a second computational cycle
                                                                            of the numerical simulation.
        tional cycle of the numerical simulation.                  40
   10. The accelerator system as claimed in claim 9, wherein             17. The method as claimed inclaim 16, further comprising:
the accelerator controller is configured to dictate an order of          sending, by the accelerator controller, instructions for per
execution of instructions to the at least one graphics process              forming the at least one calculation to the at least one
ing unit.                                                                   graphics processing unit.
   11. The accelerator system as claimed in claim 9, wherein 45          18. The method as claimed in claim 16, further comprising:
the accelerator controller is configured to send instructions to         partitioning the accelerator memory into the first partition,
the at least one graphics processing unit, wherein the at least             the second partition, a third partition to store internal
one graphics processing unit is configured to execute the                   Variables, and a fourth partition to store data used as
instructions on the input data, and wherein the accelerator                 input at a particular computation cycle of the numerical
                                                                            simulation.
controller is configured to transfer the output data from the 50         19. The method of claim 16, further comprising:
accelerator memory to a main memory during the execution                 transferring, by the accelerator controller, the first output
of the instructions by the at least one graphics processing unit.           data to the main memory during the second computa
   12. The accelerator system as claimed in claim 9, wherein                tional cycle.
the accelerator memory comprises a texture memory bank to
Store the input data and the output data and a shader memory