jain.com
Public record. We host this document directly; the copy served here does not depend on any third party. Retrieved September 29, 2026.
Read the documentAlso at archive.org ↗

Neural AI, LLC v. Tesla Inc. — Entry #6: CORRECTED MOTION to Compel Compliance With Subpoena Served on Third Party Tesla, Inc

Case: Neural AI, LLC v. Tesla Inc. txwd · 7:26-cv-00318

filed August 17, 2026

What this document is

Docket entry #6 · filed August 18, 2026

CORRECTED MOTION to Compel Compliance With Subpoena Served on Third Party Tesla, Inc. by Neural AI, LLC. (Attachments: # 1 Affidavit Declaration of Tanner Laiche, # 2 Exhibit 1, # 3 Exhibit 2, # 4 Exhibit 3, # 5 Exhibit 4, # 6 Exhibit 5, # 7 Exhibit 6, # 8 Exhibit 7, # 9 Exhibit 8, # 10 Exhibit 9, # 11 Exhibit 10, # 12 Exhibit 11, # 13 Exhibit 12, # 14 Exhibit 13, # 15 Exhibit 14, # 16 Exhibit 15, # 17 Exhibit 16, # 18 Exhibit 17, # 19 Exhibit 18, # 20 Exhibit 19, # 21 Exhibit 20, # 22 Exhibit 21, # 23 Proposed Order)(Magni, Rocco) (Entered: 08/18/2026)

Who is involved

Why we have it

We follow this case because it names a company we track, although that company is not a party:

A free copy from the RECAP archive of federal court filings (mirrored at the Internet Archive), retrieved September 29, 2026. Federal court filings are public records.

URL
https://archive.org/download/gov.uscourts.txwd.1172927763/gov.uscourts.txwd.1172927763.6.4.pdf
Kind
court_filing
Publisher
RECAP
Retrieved
2026-09-29 05:58:26.026496-04:00
HTTP status
200
MIME
application/pdf
Bytes
2170190
SHA-256
bcdeba47b6d400e239511e488d8daed744851817306ea6527b9c8795109f39d0

Document text

22 page(s), 144,947 characters, converted from the PDF's text layer · plain text.

Full text
Case 7:26-mc-00318-LS   Document 6-4   Filed 08/18/26   Page 1 of 22


                EXHIBIT

                              3


        Case 7:26-mc-00318-LS                                      Document 6-4                                                    Filed 08/18/26                            Page 2 of 22

                                                                                                                                                             USOORE48438E
(( 1912)) United
          United States
          Reissued Patent                                                              ( 10 ) Patent Number :       US RE48,438 E
      Gorchetchnikov et al .                                                           (45 ) Date of Reissued Patent : * Feb . 16, 2021
( 54 ) GRAPHIC PROCESSOR BASED                                                                             ( 56 )                                            References Cited
       ACCELERATOR SYSTEM AND METHOD
                                                                                                                                                  U.S. PATENT DOCUMENTS
( 71 ) Applicant: Neurala, Inc. , Boston, MA (US)                                                                        5,063,603 A                          11/1991 Burt
( 72 ) Inventors: Anatoli Gorchetchnikov , Belmont, MA                                                                   5,136,687 A                          8/1992 Edelman et al .
                     (US ) ; Heather Marie Ames , Milton ,                                                                                                      (Continued )
                     MA (US ) ; Massimiliano Versace,                                                                                    FOREIGN PATENT DOCUMENTS
                     Milton, MA (US ); Fabrizio Santini ,
                     Jamaica Plain, MA (US)                                                                EP                    1 224 622 B1                            7/2002
                                                                                                           WO               WO 2014/190208                              11/2014
( 73 ) Assignee : Neurala, Inc. , Boston, MA (US)                                                                                                               ( Continued )
( * ) Notice:        This patent is subject to a terminal dis
                     claimer .                                                                                                                         OTHER PUBLICATIONS
( 21 ) Appl. No .: 15 /808,201                                                                             Cornwall et al . , “ Automatically Translating a General Purpose C ++
                                                                                                           Image Processing Library for GPUs ” , IEEE , Jun . 2006 , 8 pages .
(22 ) Filed :        Nov. 9 , 2017                                                                         ( Year: 2006 ) . *
                 Related U.S. Patent Documents                                                                                                                  ( Continued )
Reissue of:
( 64) Patent No .:        9,189,828                                                                       Primary Examiner William H. Wood
      Issued:             Nov. 17 , 2015                                                                  (74 ) Attorney, Agent, or Firm -Smith Baluch LLP
      Appl. No .:         14 /147,015                                                                      ( 57 )                                              ABSTRACT
       Filed :            Jan. 3 , 2014
U.S. Applications:                                                                                         An accelerator system is implemented on an expansion card
( 63 ) Continuation of application No. 11 / 860,254 , filed on                                             comprising a printed circuit board having ( a) one or more
       Sep. 24 , 2007 , now Pat . No. 8,648,867 .                                                          graphics processing units ( GPUs ) , ( b ) two or more associ
                         (Continued )                                                                      ated memory banks ( logically or physically partitioned ), (c )
                                                                                                           a specialized controller, and ( d) a local bus providing signal
( 51 ) Int. Ci.                                                                                            coupling compatible with the PCI industry standards. The
       GO6T 1/60               ( 2006.01 )                                                                 controller handles most of the primitive operations to set up
       G06F 9/50                 ( 2006.01 )                                                               and control GPU computation. Thus, the computer's central
                           (Continued )                                                                    processing unit ( CPU) can be dedicated to other tasks . In this
( 52 ) U.S. Ci .                                                                                           case a few controls ( simulation start and stop signals from
       CPC                G06T 1/20 (2013.01 ) ; G06F 9/5027
                                                                                                           the CPU and the simulation completion signal back to CPU ),
                                                                                                           GPU programs and input/output data are exchanged between
                              ( 2013.01 ) ; G06T 1/60 ( 2013.01 ) ;                                        CPU and the expansion card . Moreover, since on every time
                           ( Continued )                                                                   step of the simulation the results from the previous time step
( 58 ) Field of Classification Search                                                                      are used but not changed, the results are preferably trans
       CPC ... GO6F 9/5027 ; G06F 2209/509 ; G06T 1/20 ;                                                   ferred back to CPU in parallel with the computation .
                  GO6T 1/60 ; G06N 99/005 ; GO6N 37063
       See application file for complete search history .                                                                 56 Claims , 5 Drawing Sheets
                                                            Expansion Cards                                40

                                                                                                           420
                                                                                         ?       e.com

                                                                           403                            435
                                                                                             xxtes
                                                                                       19x VI 9XY
                                                                                             SKM 2
                                                                                                          450
                                                                                        Swarum
                                                                                       shokolaty
                                                                                        bars CPU


                                                                           19:43             Shader              GPU .

                                                                                 180
                                                                                        Outdoor
                                                                                       yoles  exi
                                                                                        Teney bank
                                                                                                                          440
                                                                                 482                                        Now external
                                                                                                                          texmex Tom FRAM
                                                                                                                                      UM
                                                                                                                           WOX + vary but


                                                        Wait for SWRP                    Svima        i                         wait for swap
                                                       otrputloutpan
                                                       xure points
                                                                                         ? output
                                                                                       texture packs
                                                                                                                                otinescioutput
                                                                                                                            extere gointate
                                                       plexulante
                                                                                                                                                       478
                                               ******** *      herbimity
                                                            Dank    AM                 leration                                                  harasana


                                                               en?
                                                                                                          490
                                                                                       Waxa ste
                                                                                         tra
                                                                                       axexkoren
                                                                                                          499


          Case 7:26-mc-00318-LS                         Document 6-4                  Filed 08/18/26                Page 3 of 22


                                                             US RE48,438 E
                                                                    Page 2

                 Related U.S. Application Data                               2014/0032461 Al        1/2014 Weng
                                                                             2014/0089232 A1        3/2014 Buibas et al .
( 60 ) Provisional application No. 60 /826,892 , filed on Sep.               2014/0052679 Al
                                                                             2015/0127149 Al
                                                                                                   11/2014 Sinyavskiy et al .
                                                                                                    5/2015 Sinyavskiy et al .
         25 , 2006 .                                                         2015/0134232 Al        5/2015 Robinson
                                                                             2015/0224648 A1        8/2015 Lee et al .
( 51 ) Int . Ci .                                                            2016/0075017 A1        3/2016 Laurent et al.
         GOOT 1/20                 ( 2006.01 )                               2016/0082597 A1        3/2016 Gorchetchnikov et al .
         GOON 37063                (2006.01 )                                2016/0096270 A1        4/2016 Gabardos et al .
         GOON 20/00                                                          2016/0198000 Al        7/2016 Gorchetchnikov et al .
                                   ( 2019.01 )                               2017/0024877 A1        1/2017 Versace et al .
( 52) U.S. Ci.                                                               2017/0076194 A1        3/2017 Versace et al .
      CPC                G06F 2209/509 (2013.01 ) ; GO6N 37063               2017/0193298 Al        7/2017 Versace et al .
                              (2013.01 ) ; GOON 20/00 (2019.01 )                        FOREIGN PATENT DOCUMENTS
( 56)                     References Cited                               WO          WO 2014/204615            12/2014
                    U.S. PATENT DOCUMENTS                                WO          WO 2015/143173             9/2015
                                                                         WO          WO 2016/014137             1/2016
        5,142,665 A * 8/1992 Bigus                          GOON 3/04
                                                               706/16                        OTHER PUBLICATIONS
        5,172,253 A 12/1992 Lynne
        5,388,206 A     2/1995 Poulton et al .                           Montrym et al . , “ The GeForce 6800 ” , IEEE , 2005 , 11 pages. ( Year :
        6,018,696 A     1/2000 Matsuoka et al .                          2005) . *
        6,336,051 B1    1/2002 Pangels et al.                            Perumalla , Kalyan S. , “ Discrete - event Execution Alternatives on
        6,647,508 B2 11/2003 Zalewski et al .
        7,119,810 B2 * 10/2006 Sumanaweera             G06T 15/005       General Purpose Graphical Processing Units (GPGPUs)", IEEE ,
                                                              345/506    “ Proceedings of the 20th Workshop on Principles of Advanced and
        7,219,085 B2 * 5/2007 Buck                          GOON 3/08    Distributed Simulation ( PADS'06 ) ” , Mar. 2006 , 8 pages . ( Year:
                                                               706/12    2006 ) . *
        7,477,256 B1 *     1/2009 Johnson                   G06F 3/14    Setoain et al . , “ Parallel Hyperspectral Image Processing on Com
                                                              345/501    modity Graphics Hardware ” , IEEE , “ Proceedings of the 2006
        7,525,547 B1 * 4/2009 Diard                         G06T 1/20    International Conference on Parallel Processing Workshops
                                                              345/502
        7,765,029 B2 7/2010 Fleischer et al .                            ( ICPPW'06 ) " , Mar. 2006 , 8 pages . ( Year: 2006 ) . *
        7,861,060 B1 * 12/2010 Nickolls et al .                712/22    Luo et al ., “ Artificial Neural Network Computation on Graphic
        7,873,650 B1       1/2011 Chapman et                             Process Unit ” , IEEE , “ Proceedings of International Joint Confer
        8,392,346 B2       3/2013 Ueda et al .                           ence on Neural Networks, Montreal, Canada, Jul. 31 - Aug. 4 , 2005 ” ,
        8,510,244 B2       8/2013 Carson et al .                         Feb. 2005 , pp . 622-626 ( Year: 2005 ) . *
        8,583,286 B2 11/2013 Fleischer et al .                           Non - Final Office Action dated Jan. 4 , 2018 from U.S. Appl. No.
        8,648,867 B2 * 2/2014 Gorchetchnikov et al. .. 345/501           15/262,637 , 23 pages.
        9,031,692 B2   5/2015 Zhu                                        Wu, Yan & J. Cai , H. ( 2010 ) . A Simulation Study of Deep Belief
        9,177,246 B2 11/2015 Buibas et al .
     9,626,566 B2          4/2017 Versace et al.                         Network Combined with the Self -Organizing Mechanism of Adap
    10,083,523 B2          9/2018 Versace et al.                         tive Resonance Theory. 10.1109 /CISE.2010.5677265 , 4 pages .
 2001/0010034 A1 *         7/2001 Burton              G06F 17/5022       Al-Kaysi, A. M. et al . , A Multichannel Deep Belief Network for the
                                                            703/17       Classification of EEG Data, from Ontology -based Information
 2002/0046271 A1           4/2002 Huang                                  Extraction for Residential Land Use Suitability: A Case Study of the
 2002/0050518 A1           5/2002 Roustaei                               City of Regina, Canada, DOI 10.1007/978-3-319-26561-2_5 , 8
 2002/0064314 Al           5/2002 Comaniciu et al .                      pages (Nov. 2015 ) .
 2002/0168100 A1 11/2002 Woodall
 2003/0026588 A1   2/2003 Elder et al .                                  Apolloni, B. et al . , Training a network of mobile neurons, Proceed
 2003/0078754 Al   4/2003 Hamza                                          ings of International Joint Conference on Neural Networks, San
 2004/0015334 Al   1/2004 Ditlow et al.                                  Jose , CA , doi: 10.1109 /IJCNN.2011.6033427, pp . 1683-1691 ( Jul.
 2005/0166042 Al   7/2005 Evans                                          31 -Aug. 5 , 2011 ) .
 2006/0129506 A1 * 6/2006 Edelman                     GO5D 1/0088        Boddapati, V., Classifying Environmental Sounds with Image Net
                                                           706/12        works, Thesis , Faculty of Computing Blekinge Institute of Tech
 2006/0184273 A1           8/2006 Sawada et al.                          nology, 37 pages . ( Feb. 2017 ) .
 2007/0052713 A1           3/2007 Chung et al .
 2007/0198222 A1           8/2007 Schuster et al .                       Khaligh -Razavi, S.-M. et al . , Deep Supervised, but Not Unsuper
 2007/0279429 A1          12/2007 Ganzer                                 vised, Models May Explain IT Cortical Represen on, PLoS
 2008/0033897 A1           2/2008 Lloyd                                  Computational Biology , vol . 10 , Issue 11 , 29 pages . (Nov. 2014 ) .
 2008/0066065 Al           3/2008 Kim et al .                            Kim , S. , Novel approaches to clustering , biclustering and algo
 2008/0258880 A1          10/2008 Smith et al .                          rithms based on adaptive resonance theory and intelligent control,
 2009/0080695 Al           3/2009 Yang et al .                           Doctoral Dissertations, Missouri University of Science and Tech
 2009/0089030 Al           4/2009 Sturrock et al .
 2009/0116688 A1           5/2009 Monacos et al .                        nology, 125 pages . ( 2016 ) .
 2010/0048242 A1           2/2010 Rhoads et al .                         Notice of Alllowance dated May 22 , 2018 from U.S. Appl. No.
 2011/0004341 Al           1/2011 Sarvadevabhatla et al .                15/262,637 , 6 pages.
 2011/0173015 A1           7/2011 Chapman et al .                        Ren, Y. et al . , Ensemble Classification and Regression Recent
 2011/0279682 A1          11/2011 Li et al .                             Developments, Applications and Future Directions , in IEEE Com
 2012/0072215 A1           3/2012 Yu et al .                             putational Intelligence Magazine , 10.1109 /MCI.2015.2471235 , 14
 2012/0089552 A1           4/2012 Chang et al .                          pages (2016) .
 2012/0197596 A1           8/2012 Comi
 2012/0316786 A1          12/2012 Liu et al .                            Sun , Z. et al . , Recognition of SAR target based on multilayer
 2013/0126703 A1           5/2013 Caulfield                              auto -encoder and SNN , International Journal of Innovative Com
 2013/0131985 Al           5/2013 Weiland et al .                        puting, Information and Control, vol . 9 , No. 11 , pp . 4331-4341 ,
 2014/0019392 A1           1/2014 Buibas et al.                          Nov. 2013 .


          Case 7:26-mc-00318-LS                              Document 6-4                  Filed 08/18/26                 Page 4 of 22


                                                                US RE48,438 E
                                                                         Page 3

( 56 )                       References Cited                                 Coifman , R.R. , Lafon , S. , Lee , A.B. , Maggioni , M. , Nadler, B. ,
                                                                              Warner, F. , and Zucker, S.W. Geometric diffusions as a tool for
                     OTHER PUBLICATIONS                                       harmonic analysis and structure definition of data : Diffusion maps.
                                                                              Proceedings of the National Academy of Sciences of the United
Adelson , E. H. , Anderson, C. H. , Bergen , J. R. , Burt, P. J. , & Ogden,   States of America , 102 ( 21 ) : 7426 , 2005 .
J. M. ( 1984 ) . Pyramid methods in image processing. RCA engineer,           Davis, C. E. 2005. Graphic Processing Unit Computation of Neural
29(6) , 33-41 .                                                               Networks. Master's thesis, University of New Mexico , Albuquer
Aggarwal, Charu C , Hinneburg, Alexander, and Keim , Daniel A. On             que, NM , 121 pages .
the surprising behavior of distance metrics in high dimensional               Dosher, B.A. , and Lu , Z.L. ( 2010 ) . Mechanisms of perceptual
space . Springer, 2001 .                                                      attention in precuing of location . Vision Res . , 40 ( 10-12 ) . 1269
Ames, H , Versace , M. , Gorchetchnikov, A. , Chandler, B. , Livitz , G. ,    1292 .
Léveillé , J. , Mingolla, E. , Carter, D. , Abdalla, H. , and Snider, G.      Ellias , S. A. , and Grossberg, S. ( 1975 ) . Pattern formation, contrast
( 2012 ) Persuading computers to act more like brains. In Advances            control and oscillations in the short term memory of shunting
in Neuromorphic Memristor Science and Applications, Kozma ,                   on -center off - surround networks. Biol Cybern 20 , pp . 69-98 .
R.Pino , R . , and Pazienza , G. ( eds ) , Springer Verlag .                  Extended European Search Report and Written Opinion dated Jun .
Ames, H. Mingolla, E. , Sohail , A. , Chandler, B. , Gorchetchnikov,          1 , 2017 from European Application No. 14813864.7 , 10 pages .
A. , Léveillé , J. , Livitz , G. and Versace, M. ( 2012 ) The Animat . IEEE   Extended European Search Report and Written Opinion dated Oct.
Pulse , Feb. 2012 , 3 ( 1 ) , 47-50 .                                         12 , 2017 from European Application No. 14800348.6 , 12 pages .
Artificial Intelligence As a Service. Invited talk, Defrag, Broomfield ,      Extended European Search Report and Written Opinion Oct. 23 ,
CO , Nov. 4-6 ( 2013 ).                                                       2017 from European Application No. 15765396.5 , 8 pages .
Aryananda, L. ( 2006 ) . Attending to learn and learning to attend for        Fazl , A. , Grossberg , S. , and Mingolla , E. ( 2009 ) . View - invariant
a social robot . Humanoids 06 , pp . 618-623 .                                object category learning, recognition, and search : How spatial and
Baraldi, A. and Alpaydin , E. ( 1998 ) . Simplified ART: A new class          object attention are coordinated using surface -based attentional
of ART algorithms. International Computer Science Institute, Berke            shrouds . Cognitive Psychology 58 , 1-48 .
ley, CA , TR- 98-004 , 1998 .                                                 Földiák , P. ( 1990 ) . Forming sparse representations by local anti
Baraldi, A. and Alpaydin , E. ( 2002 ) . Constructive feedforward ART         Hebbian learning, Biological Cybernetics, vol . 64 , pp . 165-170 .
clustering networks — Part I. IEEE Transactions on Neural Net                 Friston K. , Adams R. , Perrinet L. , & Breakspear M. ( 2012 ) . Per
works 13 ( 3 ) , 645-661 .                                                    ceptions as hypotheses: saccades as experiments . Frontiers in Psy
Baraldi, A. and Parmiggiani, F. ( 1997 ) . Fuzzy combination of               chology, 3 ( 151 ), 1-20 .
Kohonen's and ART neural network models to detect statistical                 Galbraith , B.V, Guenther, F.H. , and Versace , M. ( 2015 ) A neural
regularities in a random sequence of multi - valued input patterns. In        network -based exploratory learning and motor planning system for
International Conference on Neural Networks, IEEE .                           co - robots.Frontiers in Neuroscience, in press .
Baraldi, Andrea and Alpaydin , Ethem . Constructive feedforward               George , D. and Hawkins, J. ( 2009 ) . Towards a mathematical theory
ART clustering networks — part II . IEEE Transactions on Neural               of cortical micro - circuits . PLoS Computational Biology 5 ( 10 ) , 1-26 .
Networks, 13 ( 3 ) : 662-677 , May 2002. ISSN 1045-9227 . doi: 10.1109/       Georgeii , J. , and Westermann , R. ( 2005 ) . Mass - spring systems on
tnn.2002.1000131
1000131 .
                 . URL http://dx.doi.org/10.1109/tnn.2002 .                   the GPU . Simulation Modelling Practice and Theory 13 , pp . 693
                                                                              702 .
Bengio , Y., Courville, A. , & Vincent, P. Representation learning: A         Gorchetchnikov A. , Hasselmo M.E. ( 2005 ) . A biophysical imple
review and new perspectives , IEEE Transactions on Pattern Analy              mentation of a bidirectional graph search algorithm to solve mul
sis and Machine Intelligence , vol . 35 Issue 8 , Aug. 2013 , pp .            tiple goal navigation tasks . Connection Science , 17 ( 1-2 ) , pp . 145
1798-1828 .                                                                   166 .
Berenson , D. et al ., A robot path planning framework that learns            Gorchetchnikov A. , Hasselmo M.E. ( 2005 ) . A simple rule for
from experience, 2012 International Conference on Robotics and                spike -timing -dependent plasticity : local influence of AHP current.
Automation, 2012 , 9 pages [ retrieved from the internet] URL : http : //     Neurocomputing, 65-66 , pp . 885-890 .
users.wpi.edu/-dberenson/lightning.pdf.                                       Gorchetchnikov A. , Versace, M. , Hasselmo M.E. ( 2005 ) . A Model
Bernhard , F. , and Keriven, R. ( 2005 ) . Spiking Neurons on GPUs.           of STDP Based on Spatially and Temporally Local Information :
Tech . Rep. 05-15 , Ecole Nationale des Ponts et Chauss’es, 8 pages.          Derivation and Combination with Gated Decay. Neural Networks,
Besl , P. J. , & Jain , R. C. ( 1985 ) . Three - dimensional object recog     18 , pp . 458-466 .
nition . ACM Computing Surveys (CSUR ), 17 ( 1 ) , 75-145 .                   Gorchetchnikov A. , Versace, M. , Hasselmo M.E. ( 2005 ) . Spatially
Bohn , C.-A. Kohonen . ( 1998 ) . Feature Mapping Through Graphics            and temporally local spiketiming -dependent plasticity rule. In :
Hardware. In Proceedings of 3rd Int. Conference on Computational              Proceedings of the International Joint Conference on Neural Net
Intelligence and Neurosciences , 4 pages .                                    works, No. 1568 in IEEE CD - ROM Catalog No. 05CH37662C , pp .
Bradski, G. , & Grossberg, S. ( 1995 ) . Fast -learning Viewnet archi         390-396 .
tectures for recognizing three - dimensional objects from multiple            Gorchetchnikov A. An Approach to a Biologically Realistic Simu
two - dimensional views. Neural Networks, 8 ( 7-8 ) , 1053-1080 .             lation of Natural Memory . Master's thesis , Middle Tennessee State
Canny, J.A. ( 1986 ) . Computational Approach to Edge Detection,              University, Murfreesboro, TN , 70 pages .
IEEE Trans. Pattern Analysis and Machine Intelligence , 8 (6 ):679            Grossberg, S. ( 1973 ) . Contour enhancement, short -term memory ,
698 .                                                                         and constancies in reverberating neural networks. Studies in Applied
Carpenter, G.A. and Grossberg, S. ( 1987 ) . A massively parallel             Mathematics 52 , 213-257 .
architecture for a self -organizing neural pattern recognition machine .      Grossberg , S. , and Huang, T.R. ( 2009 ) . Artscene : A neural system
Computer Vision, Graphics, and Image Processing 37 , 54-115 .                 for natural scene classification . Journal of Vision, 9 (4 ) , 6.1-19 .
Carpenter, G.A. , and Grossberg , S. ( 1995 ) . Adaptive resonance            doi : 10.1167 /9.4.6 .
theory ( ART ). In M. Arbib ( Ed . ) , The handbook of brain theory and       Grossberg, S. , and Versace , M. ( 2008 ) Spikes , synchrony, and
neural networks. (pp . 79-82 ) . Cambridge , M.A .: MIT press .               attentive learning by laminar thalamocortical circuits . Brain Research ,
Carpenter, G.A. , Grossberg, S. and Rosen , D.B. ( 1991 ) . Fuzzy ART:        1218C , 278-312 [ Authors listed alphabetically ].
Fast stable learning and categorization of analog patterns by an              Hagen , T.R. , Hjelmervik, J. , Lie , K.-A. , Natvig , J. and Ofstad
adaptive resonance system . Neural Networks 4 , 759-771 .                     Henriksen , M. ( 2005 ) . Visual simulation of shallow -water waves .
Carpenter, Gail A and Grossberg, Stephen . The art of adaptive                Simulation Modelling Practice and Theory 13 , pp . 716-726 .
pattern recognition by a self -organizing neural network . Computer,          Hasselt , Hado Van . Double q -learning. In Advances in Neural
21 ( 3 ) : 77-88 , 1988 .                                                     Information Processing Systems , pp . 2613-2621 , 2010 .
Coifman , R.R. and Maggioni , M. Diffusion wavelets . Applied and             Hinton , G. E. , Osindero , S. , and Teh , Y. ( 2006 ) . A fast learning
Computational Harmonic Analysis, 21 ( 1 ) : 53-94 , 2006 .                    algorithm for deep belief nets. Neural Computation, 18 , 1527-1554 .


          Case 7:26-mc-00318-LS                               Document 6-4                  Filed 08/18/26                   Page 5 of 22


                                                                US RE48,438 E
                                                                          Page 4

( 56 )                    References Cited                                     Léveillé , J. , Ames, H. , Chandler, B. , Gorchetchnikov, A. , Mingolla ,
                                                                               E. , Patrick , S. , and Versace, M. ( 2010 ) Learning in a distributed
                     OTHER PUBLICATIONS                                        software architecture for large -scale neural modeling. BIONET
                                                                               ICS10 , Boston , MA , USA .
Hodgkin , A.L. , and Huxley, A.F. ( 1952 ) . Quantitative description of       Livitz , G. , Versace, M. , Gorchetchnikov, A. , Vasilkoski, Z. , Ames,
membrane current and its application to conduction and excitation              H. , Chandler, B. , Léveillé , J. , Mingolla , E. , Snider, G. , Amerson , R. ,
in nerve. J Physiol 117 , pp . 500-544 .                                       Carter, D. , Abdalla, H. , and Qureshi , S. ( 2011 ) Visually -Guided
Hopfield , J. ( 1982 ) . Neural networks and physical systems with             Adaptive Robot ( ViGUAR ). Proceedings of the International Joint
emergent collective computational abilities . In Proc Natl Acad Sci            Conference on Neural Networks ( IJCNN) 2011 , San Jose , CA ,
USA , vol . 79 , pp . 2554-2558 .                                              USA .
Ilie , A. ( 2002 ) . Optical character recognition on graphics hardware .      Lowe , D.G. ( 2004 ) . Distinctive Image Features from Scale - Invariant
Tech . Rep . integrative paper, UNCCH , Department of Computer                 Keypoints. Journal International Journal of Computer Vision archive
Science, 9 pages .                                                             vol . 60 , 2 , 91-110 .
International Preliminary Report on Patentability in related PCT               Lu , Z.L. , Liu , J. , and Dosher, B.A. ( 2010 ) Modeling mechanisms of
Application No. PCT/US2014 /039162 filed May 22 , 2014 , dated                 perceptual learning with augmented Hebbian re -weighting. Vision
Nov. 24 , 2015 , 7 pages.                                                      Research , 50 ( 4 ) . 375-390 .
International Preliminary Report on Patentability in related PCT               Mahadevan, S. Proto -value functions: Developmental reinforce
Application No. PCT/US2014 /039239 filed May 22 , 2014 , dated                 ment learning. In Proceedings of the 22nd international conference
Nov. 24 , 2015 , 8 pages.                                                      on Machine learning, pp . 553-560 . ACM , 2005 .
International Preliminary Report on Patentability dated Nov. 8 ,               Meuth , J.R. and Wunsch , D.C. ( 2007 ) A Survey of Neural Compu
2016 from International Application No. PCT/US2015 /029438, 7                  tation on Graphics Processing Hardware. 22nd IEEE International
pages.                                                                         Symposium on Intelligent Control, Part of IEEE Multi-conference
International Search Report and Written Opinion dated Feb. 18 ,                on Systems and Control, Singapore, Oct. 1-3 , 2007 , 5 pages .
2015 from International Application No. PCT /US2014 /039162, 12                Mishkin M , Ungerleider LG . ( 1982 ) . " Contribution of striate inputs
pages.                                                                         to the visuospatial functions of parieto -preoccipital cortex in mon
International Search Report and Written Opinion dated Feb. 23 ,                keys,” Behav Brain Res, 6 ( 1 ) : 57-77 .
2016 from International Application No. PCT /US2015 /029438 , 11               Mnih , Volodymyr, Kavukcuoglu , Koray, Silver, David , Rusu , Andrei
pages .                                                                        A , Veness, Joel , Bellemare, Marc G , Graves, Alex , Riedmiller,
International Search Report and Written Opinion dated Jul. 6 , 2017            Martin , Fidjeland, Andreas K , Ostrovski, Georg , et al . Human -level
from International Application No. PCT/US2017 /029866 , 12 pages .             control through deep reinforcement learning. Nature, 518 (7540 ):529
International Search Report and Written Opinion dated Nov. 26 ,                533 , Feb. 25 , 2015 .
2014 from International Application No. PCT /US2014 /039239, 14                Moore , Andrew W and Atkeson, Christopher G. Prioritized sweep
pages.                                                                         ing : Reinforcement learning with less data and less time . Machine
International Search Report and Written Opinion dated Sep. 15 ,                Learning, 13 ( 1 ) : 103-130 , 1993 .
2015 from International Application No. PCT /US2015 /021492 , 9                Najemnik, J. , and Geisler, W. ( 2009 ) . Simple summation rule for
pages.                                                                         optimal fixation selection in visual search . Vision Research . 49 ,
Itti, L. , and Koch , C. ( 2001 ) . Computational modelling of visual          1286-1294 .
attention . Nature Reviews Neuroscience, 2 ( 3 ) , 194-203 .                   Notice of Allowance dated Jul. 27 , 2016 from U.S. Appl. No.
Itti, L. , Koch , C. , and Niebur, E. ( 1998 ) . A Model of Saliency - Based   14/662,657 .
Visual Attention for Rapid Scene Analysis, 1-6 .                               Notice of Allowance dated Dec. 16 , 2016 from U.S. Appl. No.
Jarrett, K. , Kavukcuoglu , K. , Ranzato , M. A. , & LeCun , Y. ( Sep.         14/662,657 .
2009 ) . What is the best multi - stage architecture for object recog          Oh , K.-S. , and Jung, K. ( 2004 ) . GPU implementation of neural
nition ?. In Computer Vision , 2009 IEEE 12th International Confer             networks. Pattern Recognition 37 , pp . 1311-1314 .
ence on (pp . 2146-2153 ) . IEEE .                                             Oja, E. ( 1982 ) . Simplified neuron model as a principal component
Kipfer, P., Segal , M. , and Westermann , R. ( 2004 ) . UberFlow : A           analyzer. Journal of Mathematical Biology 15 ( 3 ) , 267-273 .
GPU - Based Particle Engine . In Proceedings of the SIGGRAPH/                  Partial Supplementary European Search Report dated Jul. 4 , 2017
Eurographics Workshop on Graphics Hardware 2004 , pp . 115-122 .               from European Application No. 14800348.6 , 13 pages .
Kolb , A. , L. Latta , and C. Rezk - Salama. ( 2004 ) . “ Hardware - Based     Raijmakers, M.E.J. , and Molenaar, P. ( 1997 ) . Exact ART: A com
Simulation and Collision Detection for Large Particle Systems.” in             plete implementation of an ART network Neural networks 10 (4 ) ,
Proceedings of the SIGGRAPH /Eurographics Workshop on Graph                    649-669 .
ics Hardware 2004 , pp . 123-131 .                                             Ranzato , M. A. , Huang, F. J. , Boureau , Y. L. , & Lecun , Y. ( Jun .
Kompella, Varun Raj, Luciw , Matthew , and Schmidhuber, Jürgen.                2007 ) . Unsupervised learning of invariant feature hierarchies with
Incremental slow feature analysis: Adaptive low - complexity slow              applications to object recognition . In Computer Vision and Pattern
feature updating from high -dimensional input streams. Neural Com              Recognition, 2007. CVPR’07 . IEEE Conference on (pp . 1-8 ) . IEEE .
putation , 24 ( 11 ) : 2994-3024 , 2012 .                                      Raudies, F. , Eldridge, S. , Joshi , A. , and Versace, M. ( Aug. 20 , 2014 ) .
Kowler, E. ( 2011 ) . Eye movements: The past 25years. Vision                  Learning to navigate in a virtual world using optic flow and stereo
Research , 51 ( 13 ) , 1457-1483 . doi : 10.1016 / j.visres.2010.12.014 .      disparity signals. Artificial Life and Robotics, DOI 10.1007/ s10015
Larochelle H. , & Hinton G. ( 2012 ) . Learning to combine foveal              014-0153-1 .
glimpses with a third - order Boltzmann machine . NIPS 2010,1243               Riesenhuber, M. , & Poggio , T. ( 1999 ) . Hierarchical models of
1251 .                                                                         object recognition in cortex . Nature Neuroscience , 2 ( 11 ) , 1019
LeCun , Y., Kavukcuoglu, K. , & Farabet, C. (May 2010 ) . Convolu              1025 .
tional networks and applications in vision . In Circuits and Systems           Riesenhuber, M. , & Poggio , T. ( 2000 ) . Models of object recogni
( ISCAS ) , Proceedings of 2010 IEEE International Symposium on                tion . Nature neuroscience , 3 , 1199-1204 .
( pp . 253-256 ) . IEEE .                                                      Rolfes, T. ( 2004 ) . Artificial Neural Networks on Programmable
Lee , D. D. and Seung, H. S. ( 1999 ) . Learning the parts of objects          Graphics Hardware . In Game Programming Gems 4 , A. Kirmse, Ed .
by non -negative matrix factorization. Nature, 401 ( 6755 ) : 788-791 .        Charles River Media , Hingham , MA , pp . 373-378 .
Lee , D. D. , and Seung , H. S. ( 1997 ) . “ Unsupervised learning by          Rublee , E. , Rabaud, V., Konolige , K. , & Bradski, G. ( 2011 ) . ORB :
convex and conic coding .” Advances in Neural Information Pro                  An efficient alternative to SIFT or SURF. In IEEE International
cessing Systems , 9 .                                                          Conference on Computer Vision ( ICCV ) 2011 , 2564-2571 .
Legenstein , R. , Wilbert, N. , and Wiskott, L. Reinforcement learning         Ruesch, J. et al . ( 2008 ) . Multimodal Saliency -Based Bottom -Up
on slow features of high - dimensional input streams. PLoS Compu               Attention a Framework for the Humanoid Robot iCub . 2008 IEEE
tational Biology, 6 ( 8 ) , 2010. ISSN 1553-734X .                             International Conference on Robotics and Automation , pp . 962-965 .


          Case 7:26-mc-00318-LS                              Document 6-4                 Filed 08/18/26                   Page 6 of 22


                                                                US RE48,438 E
                                                                         Page 5

( 56 )                    References Cited                                    Tong , F. , Ze -Nian Li , ( 1995 ) . Reciprocal -wedge transform for
                                                                              space - variant sensing, Pattern Analysis and Machine Intelligence,
                    OTHER PUBLICATIONS                                        IEEE Transactions on , vol . 17 , No. 5 , pp . 500-51 . doi: 10.1109 /
                                                                              34.391393 .
Rumelhart D. , Hinton G. , and Williams, R. ( 1986 ) . Learning inter         Torralba, A. , Oliva , A. , Castelhano, M.S. , Henderson, J.M. ( 2006 ) .
nal representations by error propagation. In Parallel distributed             Contextual guidance of eye movements and attention in real -world
processing : explorations in the microstructure of cognition , vol . 1 ,      scenes: the role of global features in object search . Psychological
MIT Press .                                                                   Review , 113 ( 4 ) .766-786 .
Rumpf, M. and Strzodka, R. Graphics processor units : New pros                Van Hasselt , Hado, Guez, Arthur, and Silver, David . Deep rein
pects for parallel computing. In Are Magnus Bruaset and Aslak                 forcement learning with double q - learning. arXiv preprint arXiv :
                                                                              1509.06461 , Sep. 22 , 2015 .
Tveito , editors , Numerical Solution of Partial Differential Equations       Versace, M. ( 2006 ) From spikes to interareal synchrony: how
on Parallel Computers, vol . 51 of Lecture Notes in Computational             attentive matching and resonance control learning and information
Science and Engineering , pp . 89-134 . Springer, 2005 .                      processing by laminar thalamocortical circuits . NSF Science of
Schaul, Tom , Quan , John , Antonoglou, Ioannis, and Silver, David .          Learning Centers PI Meeting, Washington , DC , USA .
Prioritized experience replay. arXiv preprint arXiv: 1511.05952 ,             Versace, M. , ( 2010 ) Open - source software for computational neu
Nov. 18 , 2015 .                                                              roscience: Bridging the gap between models and behavior. In
Schmidhuber, J. ( 2010 ) . Formal theory of creativity, fun , and             Horizons in Computer Science Research ,vol. 3 .
intrinsic motivation ( 1990-2010 ) . Autonomous Mental Develop                Versace, M. , Ames, H. , Léveillé, J. , Fortenberry, B. , and Gorchetchnikov,
                                                                              A. ( 2008 ) KinNeSS : A modular framework for computational
ment, IEEE Transactions on , 2 ( 3 ) , 230-247 .                              neuroscience. Neuroinformatics, 2008 Winter ; 6 (4 ) : 291-309 . Epub
Schmidhuber, Jürgen . Curious model -building control systems. In             Aug 10 , 2008 .
Neural Networks, 1991. 1991 IEEE International Joint Conference               Versace, M. , and Chandler, B. ( 2010 ) MONETA : A Mind Made from
on , pp . 1458-1463 . IEEE , 1991 .                                           Memristors. IEEE Spectrum , Dec. 2010 .
Seibert, M. , & Waxman , A.M. ( 1992 ) . Adaptive 3 - D Object Rec            Webster, Bachevalier, Ungerleider ( 1994 ) . Connections of IT areas
ognition from Multiple Views. IEEE Transactions on Pattern Analy              TEO and TE with parietal and frontal cortex in macaque monkeys.
sis and Machine Intelligence , 14 ( 2 ) , 107-124 .                           Cerebal Cortex, 4 ( 5 ) , 470-483 .
Sherbakov, L. and Versace, M. ( 2014 ) Computational principles for           Wiskott, Laurenz and Sejnowski, Terrence . Slow feature analysis:
an autonomous active vision system . Ph.D., Boston University,
                                                                              Unsupervised learning of invariances. Neural Computation , 14 (4 ):715
                                                                              770 , 2002 .
http://search.proquest.com/docview/1558856407.                                Livitz G. , Versace M. , Gorchetchnikov A. , Vasilkoski Z. , Ames H. ,
Sherbakov, L. , Livitz , G. , Sohail , A. , Gorchetchnikov, A. , Mingolla ,   Chandler B. , Leveille J. and Mingolla E. ( 2011 ) Adaptive, brain -like
E. , Ames, H. , and Versace, M. ( 2013a) CogEye: An online active             systems give robots complex behaviors, The Neuromorphic Engi
vision system that disambiguates and recognizes objects. NeuComp              neer, : 10.2417 / 1201101.003500 Feb. 2011. 3 pages .
2013 .                                                                        Salakhutdinov, R. , & Hinton, G. E. ( 2009 ) . Deep boltzmann machines.
Sherbakov, L. , Livitz , G. , Sohail , A. , Gorchetchnikov, A. , Mingolla ,   In International Conference on Artificial Intelligence and Statistics
E. , Ames, H. , and Versace, M ( 2013b) A computational model of the          (pp . 448-455 ).
role of eye -movements in object disambiguation . Cosyne, Feb.                Sherbakov, L. et al . 2012. CogEye : from active vision to context
28 -Mar. 3 , 2013. Salt Lake City, UT, USA .                                  identification, youtube, retrieved from the Internet on Oct. 10 , 2017 :
Smolensky, P. ( 1986 ) . Information processing in dynamical sys              URL : //www.youtube.com/watch ? v = i5PQk962B1k, 1 page .
tems : Foundations of harmony theory. In D. E.                                Sherbakov, L. et al. 2013. CogEye: system diagram module brain
Spratling, M. W. ( 2008 ) . Predictive coding as a model of biased            area function algorithm approx # neurons, retrieved from the
competition in visual attention . Vision Research, 48 ( 12 ) : 1391-1408 .    Internet on Oct. 12 , 2017 : URL : //http ://www-labsticc.univ-ubs.
Spratling, M. W. ( 2012 ) . Unsupervised learning of generative and           fr / ~ coussy /neucomp2013 / index_fichiers /material / posters /
discriminative weights encoding elementary image components in a              NeuComp2013_final56x36.pdf, 1 page .
predictive coding model of cortical function . Neural Computation ,           Snider, Greg, et al . “ From synapses to circuitry : Using memristive
24 ( 1 ) : 60-103 .                                                           memory to explore the electronic brain . ” IEEE computer, vol . 44 ( 2 ) .
Spratling, M. W. , De Meyer, K. , and Kompass , R. ( 2009 ) . Unsu            (2011 ) : 21-28 .
pervised learning of overlapping image components using divisive              Versace, TEDx Fulbright, Invited talk, Washington DC , Apr. 5 ,
input modulation . Computational intelligence and neuroscience .              2014. 30 pages .
Sprekeler, H. On the relation of slow feature analysis and laplacian          Versace , Brain - inspired computing. Invited keynote address, Bionet
eigenmaps. Neural Computation, pp . 1-16 , 2011 .                             ics 2010 , Boston , MA , USA . 1 page .
Sutton , Richard S and Barto , Andrew G. Reinforcement learning:
An introduction. MIT Press , 1998 .                                           * cited by examiner


   Case 7:26-mc-00318-LS   Document 6-4        Filed 08/18/26   Page 7 of 22


U.S. Patent       Feb.i6, 2021        Sheet 1 of 5              US RE48,438 E


                                                                   14? *


                                 130
                       ????????????


   Case 7:26-mc-00318-LS     Document 6-4    Filed 08/18/26   Page 8 of 22


U.S. Patent       Feb. 16 , 2021    Sheet 2 of 5              US RE48,438 E


    I


                      220


                                         I


                                                         RAM


       Case 7:26-mc-00318-LS                                                                                               Document 6-4                                                   Filed 08/18/26                                        Page 9 of 22


U.S. Patent                                                                            Feb. 16 , 2021                                                              Sheet 3 of 5                                                                 US RE48,438 E


                                                                                        13:25                                        326                                                              330                                                                    FIG
                                                                                                                                                                                                                                                                             3
                                                                                                                                                                                                                                                                             .
                                320


               CEXPANRSIDO CONTRLE INTALZ O
                   180
                                                                                           IEXNTPRUATL FTEXRUOEMS TOTREXATUMRE BMEANORKY POULATIN BSIHNADRES RAMTOFROM MSEHAODRY BANK                                      COMPUTAIN   )
                                                                                                                                                                                                                                       4
                                                                                                                                                                                                                                       .
                                                                                                                                                                                                                                       FIG
                                                                                                                                                                                                                                       SEE
                                                                                                                                                                                                                                       (


                                COMPUTAINL   3ST0RE3AM
                                                                                         309
                                                                                                   DATA
                                                                                                                                   310
                                                                                                                                                     PSDAHRSDTEAR AGENRDTO
                                                                                                                                                                                        316
                                                                                                                                                                                              BATA
                                                                                                                                                                                                               314
                                                                                                                                                                                                                          DATA
                                                                                                                                                                                                                                                                       350

                   304
                                                                                                PIANRSUETR TEXURE   GENRATO POULATIN COMPILER PIANRSUETR GTENXRAUOE DOUATPUAT ACUMLTION                                                 INRAM
                                                                                                                                                                                                                                                       315

                                                                                         306                                                                                                                                                              LAST ?TERATION YES
                                                                                                                                                   311                                  312

    300
                                UINTSERACION 3ST0RE2AM    307                                     IUGNRTSAEPFHIRC 305                                            SIMULATON ITALZON PROGES MONITR RESULT DISPLAY                                          NO
                                                                                                                                                                                                                                                              390
           START                                                UGRSAPEHIRC INTERFAC   INTALZO INTERACON                         USER
                                                                                                                                                                                                                          313
                                                                                                                                                                                                                                                        END


                                             308                                                                                                                                                317                                             318
  CPU120
                   ODUATPUAT 3ST0RE1AM
                                                         O
                                                         /
                                                         I
                                                         DISK


                                                                INTALZO                                                                                                                                DOUATPUAT TODISK
                                                                                                                                                                                                                                                  NO
                                                                                                                                                                                                                                                      LAST I?TERATION YES


   Case 7:26-mc-00318-LS    Document 6-4             Filed 08/18/26   Page 10 of 22


U.S. Patent       Feb. 16 , 2021             Sheet 4 of 5             US RE48,438 E


         Expansion Card 180


                                   Isoss foxxos ir


   Case 7:26-mc-00318-LS     Document 6-4                                      Filed 08/18/26   Page 11 of 22


U.S. Patent        Feb. 16 , 2021                                 Sheet 5 of 5                  US RE48,438 E


                                    thepslBiofdcfw.oamhnocpxklteuniowdarhstg                    5
                                                                                                .
                                                                                                FIG


               2


       Case 7:26-mc-00318-LS                       Document 6-4               Filed 08/18/26              Page 12 of 22


                                                       US RE48,438 E
                               1                                                                    2
         GRAPHIC PROCESSOR BASED                            for input /output should be designed so that it provides the
      ACCELERATOR SYSTEM AND METHOD                         synchronization with computation .
                                                               In the case of GPGPU , the computation itself is performed
                                                            outside
Matter enclosed in heavy brackets [ ] appears in the 5 “ peripheral  of the CPU , so the complete system comprises three
original patent but forms no part of this reissue specifica hardware, and” components: user interactive hardware , disk
tion ; matter printed in italics indicates the additions ing unit ( CPUcomputational
                                                                            ) establishes
                                                                                           hardware. The central process
                                                                                          communication  and synchroni
made by reissue ; a claim printed with strikethrough zation between peripherals. Each of the peripherals           is pref
indicates that the claim was canceled, disclaimed, or held erably controlled by a dedicated thread that is executed in
invalid by a prior post- patent action or proceeding . 10 parallel with minimal interactions and dependencies on the
                                                                     other threads .
                RELATED APPLICATIONS                                  A GPU on a conventional video card is usually controlled
                                                                   through OpenGL , DirectX , or similar graphic application
   The present application is a broadening reissue applica programming
tion of U.S. Pat. No. 9,189,828, filed Jan. 3 , 2014, which 15 context of graphic interfaces (APIs ). Such APIs establish the
claims a priority benefit, under 35 U.S.C. $ 120 , as a con GPU are made . This        operations, within which all calls to the
tinuation of U.S. application Ser. No. 11 / 860,254 , now U.S. within the same threadcontext
                                                                                           of
                                                                                                   only works when initialized
                                                                                              execution that uses it . As a result,
Pat . No. 8,648,867 B2 , filed Sep. 24 , 2007 , entitled “ Graphic in a preferred embodiment, the context   is initialized within
Processor Based Accelerator System and Method , ” which in
turn claims the priority benefit, under 35 U.S.C. 8119 (e ) , of 20 aevercomputational    thread. This creates complications , how
                                                                           , in the interaction between the user interface thread that
U.S. Application No. 60/ 826,892 , filed Sep. 25 , 2006. Each
of the above - identified applications is incorporated herein by changes parameters of simulations and the computational
reference in its entirety. More than one reissue application thread that uses these parameters.
has been filed for the reissue of U.S. Pat. No. 9,189,828 ,         A solution as proposed here is an implementation of the
including this application and a reissue continuation appli- 25 computational stream of execution in hardware, so that
cation filed Dec. 29, 2020.                                      thread and context initialization are replaced by hardware
                                                                 initialization . This hardware implementation includes an
                        BACKGROUND                               expansion card comprising a printed circuit board having ( a )
                                                               one or more graphics processing units , (b ) two or more
  Graphics Processing Units (GPUs) are found in video 30 associated memory banks that are logically or physically
adapters ( graphic cards) of most personal computers ( PCs ) , partitioned, (c ) a specialized controller, and (d) a local bus
video game consoles , workstations, etc. and are considered providing signal coupling compatible with the PCI industry
highly parallel processors dedicated to fast computation of standards ( this includes but is not limited to PCI -Express ,
graphical content. With the advances of the computer and PCI - X , USB 2.0 , or functionally similar technologies ). The
console gaming industries, the need for efficient manipula- 35 controller handles most of the primitive operations needed to
tion and display of 3D graphics has accelerated the devel- set up and control GPU computation . As a result, the CPU
opment of GPUs .                                               is freed from this function and is dedicated to other tasks . In
   In addition , manufacturers of GPUs have included general this case a few controls ( simulation start and stop signals
purpose programmability into the GPU architecture leading from the CPU and the simulation completion signal back to
to the increased popularity of using GPUs for highly paral- 40 CPU) , GPU programs and input/output data are the infor
lelizable and computationally expensive algorithms outside mation exchanged between CPU and the expansion card .
of the computer graphics domain . When implemented on Moreover, since on every time step of the simulation the
conventional video card architectures, these general purpose results from the previous time step are used but not changed ,
GPU ( GPGPU) applications are not able to achieve optimal the results are preferably transferred back to CPU in parallel
performance , however. There is overhead for graphics- 45 with the computation .
related features and algorithms that are not necessary for        In general, according to one aspect , the invention features
these non-video applications.                                   a computer system . This system comprises a central pro
                                                                cessing unit , main memory accessed by the central process
                       SUMMARY                                  ing unit , and a video system for driving a video monitor in
                                                             50 response to the central processing unit as is common . The
   Numerical simulations, e.g. , finite element analysis , of computer system further comprises an accelerator that uses
large systems of similar elements ( e.g. neural networks, input data from and provides output data to the central
genetic algorithms, particle systems, mechanical systems) processing unit. This accelerator comprises at least one
are one example of an application that can benefit from graphics processing unit, accelerator memory for the graphic
GPGPU computation . During numerical simulations, disk 55 processing unit, and an accelerator controller that moves the
and user input /output can be performed independently of input data into the at least one graphics processing unit and
computation because these two processes require interac- the accelerator memory to generate the output data .
tions with peripheral hardware ( disk , screen , keyboard ,        In the preferred , the central processing unit transfers the
mouse , etc ) and put relatively low load on the central input data for a simulation to the accelerator, after which the
processing unit/system (CPU) . Complete independence is 60 accelerator executes simulation computations to generate
not desirable , however; user input might affect how the the output data, which is transferred to the central processing
computation is performed and even interrupt it if necessary. unit. Preferably, the accelerator controller dictates an order
Furthermore, the user output and the disk output are depen- of execution of instructions to the at least one graphics
dent on the results of the computation. A reasonable solution processing unit . The use of the separate controller enables
would be to separate input/output into threads, so that it is 65 data transfer during execution such that the accelerator
interacting with hardware occurs in parallel with the com- controller transfers output data from the accelerator memory
putation . In this case whatever CPU processing is required to main memory of the central processing unit .


       Case 7:26-mc-00318-LS                       Document 6-4               Filed 08/18/26             Page 13 of 22


                                                      US RE48,438 E
                               3                                                                    4
  In the preferred embodiment, the accelerator controller            not limited to , workstations, server computers, supercom
comprises an interface controller that enables the accelerator       puters, notebook computers, hand -held electronic devices
to communicate over a bus of the computer system with the such as cell phones, mp3 players, or personal digital assis
central processing unit.                                            tants ( PDAs ) , multiprocessor systems, programmable con
   In general according to another aspect , the invention also 5 sumer electronics, networks of any of the above -mentioned
features an accelerator system for a computer system , which computing devices, and distributed computing environments
comprises at least one graphics processing unit , accelerator that including any of the above -mentioned computing
memory for the graphic processing unit and an accelerator devices.
controller for moving data between the at least one graphics 10 In one implementation the GPU accelerator is imple
processing unit and the accelerator memory .
   In general according to another aspect , the invention also mented        as an expansion card 180 includes connections with
features a method for performing numerical simulations in a are installed along110with
                                                                    the motherboard        , on which the one or more CPU's 120
                                                                                                main , or system memory 130 and
computer system . This method comprises a central process mass / non volatile data storage              140 , such as hard drive or
ing unit loading input data into an accelerator system from redundant array of independent drives              (RAID ) array, for the
main memory of the central processing unit and an accel- 15 computer system 100. In the current example
erator controller transferring the input data to a graphics card 180 communicates to the motherboard ,110             the expansion
processing unit with instructions to be performed on the bus 190. This local bus 190 could be PCI , PCIviaExpress             a local
input data . The accelerator controller then transfers output PCI - X , or any other functionally similar technology (de,
data generated by the graphic processing unit to the central 20 pending upon the availability on the motherboard 110 ) . An
processing unit as output data .
   The above and other features of the invention including external version GPU accelerator is also a possible imple
various novel details of construction and combinations of mentation . In this example, the external GPU accelerator is
parts, and other advantages, will now be more particularly connected to the motherboard 110 through USB - 2.0 , IEEE
described with reference to the accompanying drawings and 1394 (Firewire ), or similar external /peripheral device inter
pointed out in the claims . It will be understood that the 25 face .
particular method and device embodying the invention are               The CPU 120 and the system memory 130 on the moth
shown by way of illustration and not as a limitation of the erboard 110 and the mass data storage system 140 are
invention . The principles and features of this invention may preferably independent of the expansion card 180 and only
be employed in various and numerous embodiments without communicate with each other and the expansion card 180
departing from the scope of the invention .                      30 through the system bus 200 located in the motherboard 110 .
                                                                    A system bus 200 in current generations of computers have
       BRIEF DESCRIPTION OF THE DRAWINGS                            bandwidths from 3.2 GB / s (Pentium 4 with AGTL + , Athlon
                                                                    XP with EVO ) to around 15 GB / s (Xeon Woodcrest with
   In the accompanying drawings, reference characters refer AGTL + , Athlon 64 /Opteron with Hypertransport), while the
to the same parts throughout the different views. The draw- 35 local bus has maximal peak data transfer rates of 4 GB / s
ings are not necessarily to scale ; emphasis has instead been (PCI Express 16 ) or 2 GB / s ( PCI -X 2.0 ) . Thus the local bus
placed upon illustrating the principles of the invention . Of 190 becomes a bottleneck in the information exchange
the drawings:                                                       between the system bus 200 and the expansion card 180. The
   FIG . 1 is a schematic diagram illustrating a computer design of the expansion card and methods proposed herein
system including the GPU accelerator according to an 40 minimizes the data transfer through the local bus 190 to
embodiment of the present invention ;                               reduce the effect of this bottleneck .
  FIG . 2 is block diagram illustrating the architecture for the       The system memory 130 is referred to as the main
GPU accelerator according to an embodiment of the present random - access memory (RAM ) in the description herein .
invention;                                                          However, this is not intended to limit the system memory
   FIG . 3 is a block / flow diagram illustrating an exemplary 45 130 to only RAM technology. Other possible computer
implementation of the top level control of the GPU accel- storage media include , but are not limited to ROM ,
erator system ;                                                     EEPROM , flash memory, or any other memory technology.
   FIG . 4 is a flow diagram illustrating an exemplary imple-          In the illustrated example, the GPU accelerator system is
mentation of the bottom level control of the GPU accelerator         implemented on an expansion card 180 on which the one or
system that is used to execute the target computation ; and 50 more GPU's 240 are mounted . It should be noted that the
  FIG . 5 is an example population of nine computational GPU accelerator system GPU 240 is separate from and
elements arranged in a 3x3 square and a potential packing independent of any GPU on the standard video card 150 or
scheme for texture pixels , according to an implementation of   other video driving hardware such as integrated graphics
the present invention .                                         systems. Thus the computations performed on the expansion
                                                             55 card 180 do not interfere with graphics display ( including
                DETAILED DESCRIPTION                            but not limited to manipulation and rendering of images ) .
                                                                  Various brand of GPU are relevant. Under current tech
   FIG . 1 shows a computer system 100 that has been nology, GPU's based on the GeForce series from NVIDIA
constructed according to the principles of the present inven- Corporation or the Catalyst series from ATI/ Advanced
tion .                                                       60 Micro Devices, Inc.
   In more detail, the computer system 100 in one example         The output to a video monitor 170 is preferably through
is a standard personal computer ( PC ) . However, this only the video card 150 and not the GPU accelerator system 180 .
serves as an example environment as computing environ- The video card 150 is dedicated to the transfer of graphical
ment 100 does not necessarily depend on or require any information and connects to the motherboard 110 through a
combination of the components that are illustrated and 65 local bus 160 that is sometimes physically separate from the
described herein . In fact, there are many other suitable            local bus 190 that connects the expansion card 180 to the
computing environments for this invention, including, but            motherboard 110 .


        Case 7:26-mc-00318-LS                    Document 6-4              Filed 08/18/26            Page 14 of 22


                                                     US RE48,438 E
                             5                                                                  6
  FIG . 2 is a block diagram illustrating the general archi-      not require hardware implementation. Also the partitioning
tecture of the GPU accelerator system and specifically the scheme is also altered based on new designs or needs of the
expansion card 180 in which at least one GPU 240 and algorithms being employed. The reason for this partitioning
associated memories 210 and 250 are mounted . Electrical is further explained in the Data Organization section, below .
( signal) and mechanical coupling with a local bus 190 5 A local bus interface 230 on the controller 220 serves as
provides signal coupling compatible with the PCI industry a driver that allows the controller 220 to communicate
standards ( this includes but is not limited to PCI , PCI -X , PCI through the local bus 190 with the system bus 200 and thus
Express, or functionally similar technology ).                     the CPU 120 and RAM 130. This local bus interface 230 is
   The GPU accelerator further preferably comprises one not intended to be limited to PCI related technology. Other
specifically designed accelerator controller 220. Depending 10 drivers can be used to interface with comparable technology
upon the implementation , the accelerator controller 220 is       as a local bus 190 .
field programmable gate array ( FPGA ) logic , or custom built       Data Organization
application - specific (ASIC ) chip mounted in the expansion         Each computational element discussed above has output
card 180 , and in mechanical and signal coupling with the variables that affect the rest of the system . For example in
GPU 240 and the associated memories 210 and 250. During 15 the case of a neural network it is the output of a neuron . A
initial design , a controller can be partially or even fully computational element also usually has several internal
implemented in software, in one example.                          variables that are used to compute output variables , but are
   The controller 220 commands the storage and retrieval of not exposed to the rest of the system , not even to other
arrays of data ( on a conventional video card the arrays of elements of the same population, typically. Each of these
data are represented as textures, hence the term ' texture' in 20 variables is represented as a texture . The important differ
this document refers to a data array unless specified other- ence between output variables and internal variables is their
wise and each element of the texture is a pixel of color access .
information ), execution of GPU programs (on a conven-               Output variables are usually accessed by any element in
tional video card these programs are called shaders , hence the system during every time step . The value of the output
the term “ shader ' in this document refers to a GPU program 25 variable that is accessed by other elements of the system
unless specified otherwise ), and data transfer between the corresponds to the value computed on the previous, not the
system bus 200 and the expansion card 180 through the local current, time step . This is realized by dedicating two textures
bus 190 which allows communication between the main to output variables — one holds the value computed during
CPU 120 , RAM 130 , and disk 140 .                                the previous time step and is accessible to all computational
   Two memory banks 210 and 250 are mounted on the 30 elements during the current time step , another is not acces
expansion card 180. In some example, these memory banks           sible to other elements and is used to accumulate new values
separated in the hardware, as shown, or alternatively imple-      for the variable computed during the current time step .
mented as a single , logically partitioned memory compo-        In -between time steps these tw te res are switched , so
nent.                                                           that newly accumulated values serve as accessible input
   The reason to separate the memory into two partitions 210 35 during the next time step , while the old input is replaced with
250 stems from the nature of the computations to which the new values of the variable. This switch is implemented by
GPU accelerator system is applied . The elements of com- swapping the address pointers to respective textures as
putation ( computational elements ) are characterized by a described in the System and Framework section .
single output variable . Such computational elements often     Internal variables are computed and used within the same
include one or more equations. Computational elements are 40 computational element. There is no chance of a race con
same or similar within a large population and are computed dition in which the value is used before it is computed or
in parallel. An example of such a population is a layer of after it has already changed on the next time step because
neurons in an artificial neural network (ANN ), where all within an element the processing is sequential. Therefore, it
neurons are described by the same equation. As a result, is possible to render the new value of internal variable into
some data and most of the algorithms are common to all 45 the same texture where the old was read from in the texture
computational elements within population , while most of the memory bank . Rendering to more than one texture from a
data and some algorithms are specific for each equation . single shader is not implemented in current GPU architec
Thus, one memory , the shader memory bank 210 , is used to tures, so computational elements that track internal variables
store the shaders needed for the execution of the required would have to have one shader per variable . These shaders
computations and the parameters that are common for all 50 can be executed in order with internal variables computed
computational elements and is coupled with the controller first, followed by output variables .
220 only. The second memory, the texture memory bank                   Further savings of texture memory is achieved through
250 , is used to store all the necessary data that are specific using multiple color components per pixel ( texture element)
for every computational element (including, but not limited to hold data . Textures can have up to four color components
to , input data , output data , intermediate results, and param- 55 that are all processed in parallel on a GPU . Thus, to
eters ) and is coupled with both the controller 220 and the maximize the use of GPU architecture it is desirable to pack
GPU 240 .                                                           the data in such a way that all four components are used by
    The texture memory bank 250 is preferably further par- the algorithm . Even though each computational element can
titioned into four sections . The first partition 250 a is have multiple variables, designating one texture pixel per
designed to hold the external input data patterns. The second 60 element is ineffective because internal variables require one
partition 250b is designed to hold the data textures repre- texture and output variables require two textures . Further
senting internal variables. The third partition 250c is more , different element types have different numbers of
designed to hold the data textures used as input at a variables and unless this number is precisely a multiple of
particular computation step on the GPU 240. The fourth four, texture memory can be wasted .
partition 250d holds the data textures used to accommodate 65 A more reasonable packing scheme would be to pack four
the output of a particular computational step on the GPU computational elements into a pixel and have separate
240. This partitioning scheme can be done logically , does textures for every variable associated with each computa


       Case 7:26-mc-00318-LS                        Document 6-4               Filed 08/18/26              Page 15 of 22


                                                       US RE48,438 E
                               7                                                                      8
tional element. In this case the packing scheme is identical             The crucial feature of the interaction between the User
for all textures, and therefore can be accessed using the same Interaction Stream 302 and the Computational Stream 303 is
algorithm . Several ways to approach this packing scheme the shift of priorities. Outside of the simulation , the system
are outlined here. An example population of nine computa- 100 is driven by the user input, thus the User Interaction
tional elements arranged in a 3x3 square (FIG . 5a ) can be 5 Stream 302 has the priority and controls the data exchange
packed by element (FIG . 5b ) , by row (FIG . 5c ) , or by square 304 between streams. After the user starts the simulation , the
( FIG . 5d) .                                                       Computational Stream 303 takes the priority and controls
   Packing by element ( FIG . 5b ) means that elements 1,2,3,4 the      data exchange between streams until the simulation is
                                                                    finished or interrupted 350 .
go into first pixel ; 5,6,7,8 go into second pixel ; 9 goes into
third pixel . This is the most compact scheme , but not 10 an The          user starts 300 the framework through the means of
convenient because the geometrical relationship is not pre theoperating           system and interacts with the software through
served during packing and its extraction depends on the size 306user         interaction section 305 of the graphic user interface
                                                                         executed on the CPU 120. The start 300 of the imple
of the population .
   Packing by row ( column; FIG . 5c ) means that elements 15 initializationbegins
                                                                    mentation            with a user action that causes a GUI
                                                                                  307 , Disk input /output initialization 308 on the
1,2,3 go into pixel ( 1,1 ) ; 3,4,5 go into pixel (2,1 ) , 7,8,9 go CPU 120 , and controller initialization 320 of the GPU
into pixel ( 3,1 ) . With this scheme the element’s y coordinate accelerator on the expansion card 180. GUI initialization
in the population is the pixel's y coordinate, while the              includes opening of the main application window and setting
element’s x coordinate in the population is the pixel's x             the interface tools that allow the user to control the frame
coordinate times four plus the index of color component. 20 work . Disk I/O initialization can be performed at the start of
Five by five populations in this case will use 2x5 texture, or the framework , or at the start of each individual simulation .
10 pixels . Five of these pixels will only use one out of four             The user interaction 305 controls the setting and editing of
components , so it wastes 37.5 % of this texture. 25x1 popu- the computational elements, parameters, and sources of
lation will use 6x1 texture ( six pixels ) and will waste 12.5 % external inputs. It specifies which equations should have
of it .                                                              25 their output saved to disk and / or displayed on the screen . It
    Packing by square ( FIG . 5d ) means that elements 1,2,4,5 allows the user to start and stop the simulation . And it
go into pixel ( 1,1 ) ; 3,6 go into pixel ( 1,2 ) ; 7,8 go into pixel performs standard interface functions such as file loading
( 2,1 ) , and 9 goes into pixel (2,2 ) . Both the row and the and saving , interactive help , general preferences and others.
column of the element are determined from the row (col-                    The user interaction 305 directs the CPU 120 to acquire
umn ) of the pixel times two plus the second ( first) bit of the 30 the new external input textures needed (this includes but is
color component index . Five by five populations in this case not limited to loading from disk 140 or receiving them in
will use 3x3 texture , or 9 pixels . Four of these pixels will real time from a recording device ), parses them if necessary
only use out of four components, and one will only use                309 , and initializes their transfer the expansion card 180 ,
one component, so it wastes 34.4 % of this texture . This is          where they are stored 325 in the texture memory bank 250
more advantageous than packing by row , since the texture is 35 by the controller 220. The user interaction 305 also directs
smaller and the waste is also lower. 25x1 population on the the CPU 120 to parse populations of elements that will be
other hand will use 13x1 texture ( thirteen pixels ) and waste used in the simulation, convert them to GPU programs
> 50 % of it , which is much worse than packing by row .       ( shaders ) , compile them 310 , and initializes their transfer to
   In order to eliminate waste altogether the population the expansion card 180 , where they are stored 326 in the
should have even dimensions in the square packing, and it 40 shader memory bank 210 by the controller 220. This opera
should have a number of columns divisible by four in row tion is accompanied by the upload 309 of the initial data into
packing. Theoretically, the chances are approximately the input partition of the texture memory bank 250 , and
equivalent for both of these cases to occur, so the particular        stores the shader order of execution in the controller 220 .
task and data sizes should determine which packing scheme The user can perform operations 309 and 310 as many times
is preferable in each individual case .                        45 as necessary prior to starting the simulation or between
   The System and Framework                                       simulations .
   FIG . 3 shows an exemplary implementation of the top             The editing of the system between simulations is difficult
level system and method that is used to control the compu- to accomplish without the hardware implementation of the
tation . It is a representation of one of several ways in which computational thread suggested herein . The system of equa
a system and method for processing numerical techniques 50 tions ( computational elements) is represented by textures
can be implemented in the invention described herein and so that track variables plus shaders that define processing
the implementation is not intended to be limited to the algorithms. As mentioned above , textures, shaders and other
following description and accompanying figure .                   graphics related constructs can only be initialized within the
   The method presented herein includes two execution rendering context, which is thread specific . Therefore tex
streams that run on the CPU 120 - User Interaction Stream 55 tures and shaders can only be initialized in the computa
302 and Data Output Stream 301. These two streams pref- tional thread .
erably do not interact directly, but depend on the same data        Network editing is a user - interactive process, which
accumulated during simulations . They can be implemented according to the scheme suggested above happens in the
as separate threads with shared memory access and executed User Interaction Stream 302. The simulation software thus
on different CPUs in the case of multi -CPU computing 60 has to take the new parameters from the User Interaction
environment. The third execution stream — Computational Stream 302 , communicate them to the Computational
Stream 303 runs on the GPU accelerator of the expansion               Stream 303 and regenerate the necessary shaders and tex
card 180 and interacts with the User Interaction Stream 302           tures . This is hard to accomplish without a hardware imple
through initialization routines and data exchange in between mentation of the Computational Stream 303. The Compu
simulations. The Computational Stream 303 interacts with 65 tational Stream 303 is forked from the User Interaction
the User Interaction Stream and the Data Output Stream Stream and it can access the memory of the parent thread,
through synchronization procedures during simulations.       but the reverse communication is harder to achieve. The


         Case 7:26-mc-00318-LS                               Document 6-4                    Filed 08/18/26                   Page 16 of 22


                                                                 US RE48,438 E
                                     9                                                                                10
controller 220 allows operations 309 and 310 to be per- timestep ( ) , TSimulator::outfileInterval ( ), and TSimulator::
formed as many times as necessary by providing the nec- outmode ( ) , the application can set the time step of the
essary communication to the User Interaction Stream 302 .       simulation , the time step of disk output, and the mode of the
   After execution of the input parser texture generation 309 5 disk output. The external input pattern should be packed into
and population parser shader generator and compiler 310 are a TPattern object and bound to the simulation object through
performed at least once , the user has the option to initialize TSimulator:: resetInputs( ) . method . TSimulator::
the simulation 311. During this initialization the main con simLength ( ) sets the length of the simulation .
trol of the framework is transferred to the GPU accelerator       The second step is to create at least one population of
system's accelerator controller 220 and computation 330 is equations       ( Tpopulation object ). Population holds one equa
started
interrupt(see
           the FIG . 4; 420, change
               simulation    ). The the
                                    userinput
                                         retains
                                             , or the abilitythe
                                                  to change  to 10 tion object TEquation. This object contains only a formula
display properties of the framework , but these interactions the           and does not hold element- specific data , so all elements of
are queued to be performed at times determined by the                            population can share single TEquation .
controller - driven data exchange 314 and 316 to avoid the before execution    The     TEquation object is converted to a GPU program
corruption of the data .                                                15                            . GPU programs have to be executed within
   The progress monitor 312 is not necessary for perfor creates this context , within
                                                                           a  graphical        context        which is stream specific . TSimulator
mance , but adds convenience. It displays the percentage of fore all programs and dataa arrays                          Computational Stream , there
completed time steps of the simulation and allows the user computation have to be initialized that                                  within
                                                                                                                                           are necessary for
                                                                                                                                              Computational
to plan the schedule using the estimates of the simulation Stream . Constructor of TPopulation is called                                          from User
wall    clock   times . Controller    - driven  data   exchange    314  20 Interaction
updates the display of the results 313. Online screen output tialized in this constructor .
                                                                                              Stream      ,  so no   GPU    - related    objects  can be ini
for the user selected population allows the user to monitor
the activity and evaluate the qualitative behavior of the to TPopulation        overcome
                                                                                                   :: fillElements ( ) is a virtual method designed
                                                                                                  this     difficulty. It is called from within the
network . Simulations with unsatisfactory behavior can be Computational Stream                                  after TSimulator       ::user
                                                                                                                                          networkCreate  ( ) is
terminated    early to  change  parameters      and  restart . Control- 25 called   in   the    User     Interaction    Stream
ler - driven data exchange 314 also drives the output of the TPopulation :: fillElements ( ) to create TEquation and other
                                                                                                                                  . A         has to override
results to disk 317. Data output to disk for convenience can computation                        related objects both element independent and
be done on an element per file basis . A suggested file format element-specific. Element independent objects include sub
includes a leftmost column that displays a simulated time for
each    of the simulation steps and subsequent columns that 30 components                       of TEquation and
                                                                           handle interdependencies                       objectsvariables
                                                                                                                     between          that describe   how to
                                                                                                                                                implemented
display variable values during this time step in all elements through                   derivatives of TGate class .
with identical equations (e.g. all neurons in a layer of a                     Element      - specific data is held in TElement objects. These
neural network ).
   Controller -driven data exchange or input parser texture objects. Therereferences
                                                                           objects      hold                      to TEquation and a set of TGate
                                                                                                   is one TElement per population, but the size
generator    316 allows the user to change input that is gen- 35 of data arrays within this object corresponds to population
erated on the fly during the simulation . This allows the size . All TElement objects have to be added to the TSimu
framework monitoring of the input that is coming from a lator list of elements by calling TSimulator:: addUnit ( )
recording device ( video camera, microphone, cell recording method                       from TPopulation :: fillElements ( ).
electrode, etc ) in real time . Similar to the initial input parser            Finally, TPopulation :: fillElements ( ) should contain a set
309
data, itarray
          preprocesses
               suitable the
                         for input
                              textureintogeneration
                                           a universaland
                                                        format   of the 40 of TElement:: add* Dependency ( ) calls for each element.
                                                              generates
textures . Unlike the initial parser 309 , here the textures are every     Each of these calls sets a corresponding dependency for
transferred to hardware not whenever ready but upon the pendentTGatepart                        object. Here TGate object holds element inde
                                                                                                            of dependency and TElement::
request of the controller 220 .                                            add * Dependency sets element-specific details.
   The controller 220 also drives the conditional testing 315 45 System provided TPopulation handles the output of com
and 318 informs the CPU - bound streams whether the simu putational                         elements, both when they need to exchange the
lation is finished . If so , the control returns to the User data and when                          they need to output it to disk . User imple
Interaction Stream . The user then can change parameters or
inputs ( 309 and 310 ) , restart the simulation (311 ) or quit the mentation                of TPopulation derivative can add screen output.
                                                                               Listing 1 is an example code of the user program that uses
framework (390 ) .                                                      50
                                                                           a  recurrent       competitive field (RCF ) equation:
   SANNDRA ( Synchronous Artificial Neuronal Network
Distributed Runtime Algorithm ; http://www.kinness.net/
Docs /SANNDRA /html) was developed to accelerate and                                                                Listing 1
optimize processing of numerical integration of large non
homogenous systems of differential equations. This library 55 uint16_t     static floatw m_compet
                                                                                          = 3 , h = 3 ; = 0.5 ;
is fully reworked in its version 2.x.x to support multiple static                 float m_persist = 1.0 ;
computational backends including those based on multicore class TCablePopRCF : public TPopulation
CPUs , GPUs and other processing systems . GPU based {TEq_RCF * m_equation ;
backend for SANNDRA -2.x.x can serve as an example            TGate * m_gatel;
practical software implementation of the method and archi- 60 TGate * m_gate2;
tecture described above and pictorially represented in FIG . void createGatingStructure( )
3.                                                           {
   To use SANNDRA, the application should create a m_gate2   m_gatel = new TGate ( 0 );
TSimulator object either directly or through inheritance . } ;        = new TGate ( 1 ) ;
This object will handle global simulation properties and 65 void createUnitStructure ( TBasicUnit* u )
control the User Interaction Stream , Data Output Stream , {
and Computational Stream . Through TSimulator ::


          Case 7:26-mc-00318-LS                                        Document 6-4         Filed 08/18/26                       Page 17 of 22


                                                                        US RE48,438 E
                                           11                                                                            12
                                     -continued                                    expensive operation , computationally, but since the textures
                                                                                   are referred to by texture IDs ( pointers ), swapping these
                                          Listing 1                                pointers for input and output textures after each time step
u-> addO20PInputDependency (m_gatel, O. , 0. , 0.004 , 0. , 0 , 0 ) ;              achieves the same result at a much lesser cost .
                                                                                 5
u- > addFullDependency (m_gate2, population ( ) );                                    In the hardware solution suggested herein , ID swapping is
}                                                                                  equivalent to swapping the base memory address for two
public : TCablePopRCF ( ) : TPopulation ( " compCPU RCF ” , w , h , true) { } ;
-TCablePopRCFO ) { if (m_equation ) delete m_equation ;                            partitions of the texture memory bank 250. They are
   if (m_gatel) delete m_gatel;
   if (m_gate2) delete m_gate2 ; } ;
                                                                                   swapped 485 during synchronization ( 485 , 430 , and 455 ) so
bool fillElements( TSimulator * sim ) ;                                         10 that data transfer 445 and the computation 435-487 proceeds
};                                                                                 immediately and in parallel with data transfer as shown in
bool TCablePopRCF :: fillElements ( TSimulatior* sim)                              FIG . 4. A hardware solution allows this parallelism through
{                                                                                  access of the controller 220 to the onboard texture memory
m_equation = new TEQ_RCF (this, m_compet, m_persist );
createGatingStructure ( ) ;                                                        bank 250 .
for( size_t i = 0 ; i < xSize ( ) ; ++ i )                                      15
                                                                                      The main computation and data exchange are executed by
 for(size_t j = 0 ; j < ySize ( ) ; ++ i)
{                                                                                  the controller 220. It runs three parallel substreams of
TElement * u = new TCPUElement(this , m_equation, i , j ) ;                        execution : Computational Substream 403 , Data Output Sub
sim-> addUnit( u );                                                                stream 402 , and Data Input Substream 404. These streams
createUnitStructure ( u );
}                                                                               20 are synchronized with each other during the swap of pointers
Return true ;                                                                      485 to the input and output texture memory partitions of the
}                                                                                  texture memory bank 250 and the check for the last iteration
int
main ( )                                                                           487. Algorithmically, these two operations are a single
{                                                                                  atomic operation, but the block diagram shows them as two
// Input pattern generation ( 309 in FIG.3 )                                    25 separate blocks for clarity.
uint32_t * pat = new uint32_t [ w * h ];
TRandom < float > randGen (0 ) ;                                                      The Computational Substream 403 performs a computa
for (uint32_t I = 0 ; I < w * h ; ++ i )
pat [i] = randGen.random ( ) ;                                                     tional cycle including a sequential execution of all shaders
Tpattern * p = new Tpattern (pat, w, h ) ;                                         that were stored in the shader memory bank 210 using the
// Setting up the simulation
                                                                              30
                                                                                   appropriate input and output textures. To begin the simula
TSimulator * cableSim = new TSimulator ( "data " ) ; // ( 308 and 320 in           tion the controller 220 initializes three execution substreams
FIG. 3)
cableSim- >timestep ( 0.05 ) ; // (320 in FIG . 3 )                              403 , 402 , and 404. On every simulation step , the Compu
cableSim-> resetInputs (p ); // (325 in FIG . 3 )                                  tational Substream 403 determines which textures the GPU
cableSim- > outfileInterval(0.1 ); // (308 in FIG . 3 )                          240 will need to perform the computations and initiates the
cableSim- > outmode (SANNDRA ::timefunc ); // ( 308 in FIG . 3 )
cableSim- > simLength (60.0 ); // (320 in FIG . 3 )                           35 upload 435 of them onto the GPU 240. The GPU 240 can
// Preparing the population                                                      communicate directly with the texture memory bank 250 to
TPopulation * cablePop new TCablePopRCFO ); // (310 in FIG . 3 )                 upload the appropriate texture to perform the computations .
cableSim- > networkCreate ( ); // (326 in FIG . 3 )                              The controller 220 also pulls the first shader (known by the
uint16_t user = 1 ;
while (user)                                                                     stored order ) from the shader memory bank 210 and uploads
{                                                                             40   450 it onto the GPU 240 .
if (! cableSim-> simulationStart ( true, 1 ) ) // ( 311 in FIG . 3 )
exit ( 1 ) ;                                                           The GPU 240 executes the following operations in this
std ::cout << " Repeat ? \ n " ; // (305 in FIG . 3 )
std :: cin >> user; // ( 305 in FIG . 3 )
                                                                    order : performs the computation ( execution of the shader )
if (user 1 )                                                        470 ; tells the controller 220 that it is done with the compu
cableSim- > networkReset ( ); // ( 305 in FIG . 3 )                 tations  for the current shader; and after all shaders for this
                                                                 45 particular equation are executed sends 480 the output tex
{
If (cableSim )                                                      tures to the output portion of the texture memory bank 250 .
Delete cableSim ; // Also deletes cablePop and its internals        This cycle continues through all of the equations based on
exit (0 ) ;
};                                                                  the branching step 482 .
                                                                 50    An example shader that performs fourth order Runge
     FIG . 4 is a detailed flow diagram illustrating a part of an Kutta numerical integration is shown in Listing 2 using
exemplary implementation of the bottom level system and GLSL notation ;
method performed during the computation on the GPU
accelerator of the expansion card 180 and is a more detailed                                      Listing 2
view of the computational box 330 in FIG . 3. FIG . 4 is a 55
representation of one of several ways in which a system and                uniform sampler2DRect Variable ;
method for processing numerical techniques can be imple                                 uniform float integration_step ;
                                                                                        float halfstep = integration_step * 0.5 ;
mented .
   With systems of equations that have complex interdepen                               float fl_6step = integration_step / 6.0 ;
                                                                                        vec4 output = texture2DRect (Variable, gl_TexCoord [0 ] .st );
dencies it is likely that the variable in some equation from 60                         // define equation here
a previous time step has to be used by some other equation                              vec4 rungekutta4 ( vec4 x )
after the new values of this variable are already computed                              {
                                                                                        const vec4 kl = equation ( x );
for new time step . To avoid data confusion , the new values                            const vec4 k2 equation ( x + halfstep * kl ) ;
of variables should be rendered in a separate texture. After                            const vec4 k3 equation ( x + halfstep * k2 );
the time step is completed for all equations, these new values 65                        const vec4 k4 = equation ( x + integration step * k3 ) ;
should be copied over old values so that they are used as                                return fl_6step * (kl + 2.0 * (k2 + k3 ) + k4 );
input during the next time step . Copying textures is an


          Case 7:26-mc-00318-LS                   Document 6-4                 Filed 08/18/26             Page 18 of 22


                                                       US RE48,438 E
                                    13                                                             14
                               -continued                                                   CONCLUSION
                                  Listing 2                            This GPU accelerator system offers the following poten
      }
                                                                    tial advantages:
      Void main (void )                                           5     1. Limited computations on the CPU 120. The CPU 120
      {                                                             is only used for user input, sending information to the
      output + = rungekutta4 (output );
      gl_FragColor = output;
                                                                    controller 220 , receiving output after each computational
      }                                                     cycle ( or less frequently as defined by the user ), writing this
                                                            output to disk 140 , and displaying this output on the monitor
                                                         10 170. This frees the CPU 120 to execute other applications
  The shader in Listing 2 can be executed on conventional and allows the expansion card to run at its full capacity
video card . Using the controller 220 this code can be further         without being slowed down by extensive interactions with
optimized , however. Since the integration step does not               the CPU 120 .
change during the simulation , the step itself as well as the            2. Minimizing data transfer between the expansion card
halfstep and % of the step can be computed once per 15 180 and the system bus 200. All of the information needed
simulation , and updated in all shaders by a shader update to perform the simulations will be stored on the expansion
procedures 310 , 326 discussed above .                            card 180 and all simulations will take place on it . Further
  After all of the equations in the computational cycle are       more , whatever data transfer remains necessary will take
computed the main execution substream 403 on the control          place in parallel with the computation , thus reducing the
ler 220 can switch 485 the reference pointers of the input and 20 impact
                                                                     3. New of this
                                                                               waytransfer   on the
                                                                                      to execute GPUperformance
                                                                                                       programs (. shaders ). Previ
output portions of the texture memory bank 250 .                  ously, the CPU 120 had full control over the order of
   The two other substreams of execution on the controller        shader's execution and was required to produce specific
220 are waiting ( blocks 430 and 455 , respectively) for this commands          on every cycle to tell the GPU 240 which shader
switch  to begin their execution . The Data Input  Substream   25 to  use .  With   the invention disclosed herein , shaders will
404 is controlling 440 the input of additional data from the initially be stored on the shader memory bank 210 on the
CPU 120. This is necessary in cases where the simulation is expansion card 180 and will be sent to the GPU 240 for
monitoring the changing input, for example input from a execution              by the general purpose controller 220 located on
video camera or other recording device in the real time . This the expansion card .
substream uploads new external input from the CPU 120 to 30 4. Multiple parallelisms . The GPU 240 is inherently
the texture memory bank 250 so it can be used by the main parallel and is well suited to perform parallel computations.
computational substream 403 on the next computational step In parallel with the GPU 240 performing the next calcula
and waits for the next iteration 475. The Data Output tion , the controller 220 is uploading the data from the
Substream 445 controls the output of simulation results to previous calculation into main memory 130. Furthermore,
the CPU 120 if requested by the user . This substream 35 the CPU 120 at the same time uses uploaded previous results
uploads the results of the previous step to the main RAM to save them onto disk 140 and to display them on the screen
130 so that the CPU 120 can save them on disk 140 or show through the system bus 200 .
them on the results display 313 and waits for the next               5. Reuse of existing and affordable technology. All hard
iteration 460 .                                                        ware used in the invention and mentioned here - in are based
  Since the Computational Substream 403 determines the 40 on currently available and reliable components. Further
timing of input 440 and output 445 data transfers, these data          advance of these components will provide straightforward
transfers are driven by the controller 220. To further reduce          improvements of the invention .
the data transfer overhead ( and disk 140 overhead also ) the            While this invention has been particularly shown and
controller 220 initiates transfer only after selected compu-           described with references to preferred embodiments thereof,
tational steps . For example, if the experimental data that is 45 it will be understood by those skilled in the art that various
simulated was recorded every 10 milliseconds (msec ) and changes in form and details may be made therein without
the simulation for better precision was computed every 1 departing from the scope of the invention encompassed by
msec , then only every tenth result has to be transferred to the appended claims .
match the experimental frequency.
   This solution stores two copies of output data , one in the 50 What is claimed is :
expansion card texture memory bank 250 and another in the         1. À computer system , comprising:
system RAM 130. The copy in the system RAM 130 is                 a central processing unit to receive input data ;
accessed twice : for disk I/O and screen visualization 313. An    main memory , operably coupled to the central processing
alternative solution would be to provide CPU 120 with a              unit via a bus , to store the input data received by the
direct read access to the onboard texture memory bank 250 55         central processing unit;
by mapping the memory of the hardware onto a global                      an accelerator, operably coupled to the central processing
memory space . The alternative solution will double the                    unit and the [first] main memory via the bus , to receive
communication through the local bus 190. Since the goal                    at least a portion of the input data from the main
discussed herein is reducing the information transfer through              memory , the accelerator comprising:
the local bus 190 , the former solution is favored .              60       at least one graphics processing unit to perform a
  The main substream 403 determines if this is the last                       sequence of computations on the at least a portion of
iteration 487. If it is the last iteration , the controller 220               the input data so as to generate output data , the
waits for the all of the execution substreams to finish 490                  sequence of computations representing an artificial
and then returns the control to the CPU 120 , otherwise it                    neural network, intermediate computations in the
begins the next computational cycle .                             65          sequence of computations representing respective
  This repeats through all of the computational cycles of the                 layers of the artificial neural network and yielding
simulation .                                                                  intermediate results; and


       Case 7:26-mc-00318-LS                         Document 6-4              Filed 08/18/26              Page 19 of 22


                                                        US RE48,438 E
                               15                                                                   16
      accelerator memory , operably coupled to the [ graphic ] system comprising a central processing unit (CPU) , a main
         at least one graphics processing unit, to store the memory operably coupled to the central processing unit via
         results of the [ plurality of sequential] sequence of a bus , an accelerator operably coupled to the CPU and the
         computations; and                                       main memory via the bus , the accelerator comprising a
   a controller, operably coupled to the at least one graphics 5 graphics processing unit (GPU) and an accelerator memory,
      processing unit and the accelerator memory, to initial- the method comprising:
      ize textures and shaders in the accelerator memory for       ( A ) performing, by the GPU , the sequence of computa
     performing the sequence of computations, to control              tions on a first portion of [ the] input data so as to
     performance of the sequence of computations by the at            generate a first portion of [the] output data , the first
      least one graphics processing unit, to transfer the at 10 portion        of the output data representing an output of a
      least a portion of the input data into the accelerator          neuron  in a first layer of the artificial neural network,
      memory during performance of the intermediate com               intermediate computations in the sequence of compu
     putations in the sequence of computations by the at              tations yielding intermediate results , wherein perform
      least one graphics processing unit, and to transfer at          ing the sequence of computations on the first portion of
      least a portion of the output data from the accelerator 15      the input data comprises ( i ) assigning an output vari
      memory to the main memory during performance of the
      intermediate computations in the sequence of compu              able to a first texture and a second texture , the output
     tations by the at least one [ graphic ] graphics processing             variable being included in a first computational ele
     unit .                                                                  ment of a plurality of computational elements, the
  2. The computer system of claim 1 , wherein the central 20                 plurality of computational elements representing the
processing unit is configured to receive the input data in                   sequence of computations and ( ii ) accumulating a first
response to a user interaction .                                             value for the output variable in the first texture during
  3. The computer system of claim 1 , wherein :                              a first time step ;
   the central processing unit is configured to receive the              ( B ) in parallel with performing the sequence of compu
      input data at a first rate ; and                              25       tations by the GPU in ( A ), transferring a second portion
   the at least one graphics processing unit is configured to                of the input data from the main memory to the accel
      perform the sequence of computations at a second rate                  erator via the bus ; [ and]
      different than the first rate .                                    ( C ) in parallel with performing the sequence of compu
   4. The computer system of claim 1 , wherein the main                      tations by the GPU in ( A ), transferring a second portion
memory is configured to store a copy of the output data 30                   of the output data from the accelerator memory to the
stored in the accelerator memory .                                           main memory via the bus, the second portion of the
   5. The computer system of claim 1 , wherein an output of                  output data representing an output of a neuron in a
at least one computation in the sequence of computations                     second layer in the artificial neural network ; and
represents an output of at least one neuron in an artificial             ( D ) performing, by the GPU , the sequence of computa
neural network .                                                    35       tions on the second portion of the input data , wherein
   6. The computer system of claim 1 , wherein accelerator                  performing the sequence of computationson the second
memory comprises:                                                           portion of the input data comprises ( i ) accumulating a
   a first memory bank to store parameters common to all of                  second value for the output variable in the second
      the computations in the sequence of computations; and                  texture during a second time step and ( ii) making the
   a second memory bank to store data specific to at least one 40           first value of the output variable in the first texture
      computation in the sequence of computations.                           accessible to other computational elements in the plu
   7. The computer system of claim 1 , wherein the controller                 rality of computational elements during the second
is configured to transfer the output data from the accelerator               time step.
memory to the main memory without transferring any of the                13. The method of claim 12 , further comprising:
intermediate results from the accelerator memory to the 45               storing the input data in the main memory in response to
main memory so as to reduce data transfer via the bus .                      a user interaction .
   8. The computer system of claim 1 , wherein the controller            14. The method of claim 12 , further comprising:
is configured to transfer at least a portion of the output data          receiving the input data at a first rate; and
from the accelerator memory to the main memory after the                 wherein ( A ) comprises performing the sequence of com
at least one graphics processing unit has begun to perform 50                putations at a second rate different than the first rate .
another sequence of computations.                                        [ 15. The method of claim 12 , wherein ( A ) comprises:
   9. The computer system of claim 8 , wherein the controller            generating an output representative of an output of at least
is configured to initiate transfer of the at least a portion of the          one neuron in an artificial neural network .]
input data and to transfer the at least a portion of the output          16. The method of claim 12 , wherein (C ) comprises:
data in parallel with performance of at least one computation 55 transferring the second portion of the output data from the
in the other sequence of computations by the at least one                   accelerator memory to the main memory without trans
graphics processing unit .                                                  ferring any of the intermediate results of the plurality of
   10. The computer system of claim 1 , wherein the con                     sequential computations from the accelerator memory
troller is configured to control execution of the sequence of               to the main memory so as to reduce data transfer via the
computations by the at least one graphics processing unit . 60              bus .
   11. The computer system of claim 1 , further comprising:              17. The method of claim 12 , wherein ( C ) comprises :
   at least one of a video camera, a microphone, or a cell               transferring the second portion of the output data from the
     recording electrode, operably coupled to the central                   accelerator memory to the main memory after the GPU
     [ processor) processing unit , to acquire the input data in       has begun to perform another sequence of computa
     real time .                                                 65   tions.
   12. A method of performing a sequence of computations            18. The method of claim 17 , wherein (C ) further com
representing an artificial neural network on a computer prises:


       Case 7:26-mc-00318-LS                      Document 6-4               Filed 08/18/26             Page 20 of 22


                                                      US RE48,438 E
                              17                                                                  18
   initiating transfer of the second portion of the output data        27. The method of claim 26 , further comprising :
      in parallel with performance of at least one computa-            storing, in a second memory partition of the memory, data
      tion in the other sequence of computations.                         specific to the first computation in the sequence of
   19. The method of claim 12 , further comprising :                      computations.
   acquiring the input data in real time with at least one of 5 28. The method of claim 27, further comprising :
      a video camera , a microphone, or a cell recording             storing, in the second memory partition , external input
      electrode operably coupled to the CPU .                            data patterns, representations of internal variables, an
   20. The method of claim 12 , further comprising :                     input of the computation in the sequence of computa
   storing parameters common to all of the computations in 10            tions , and the output of the computation in the sequence
      the sequence of computations in a first memory bank in             of computations.
      the accelerator memory ; and                                    29. The method of claim 21 , wherein storing the first
   storing data specific to at least one computation in the output data comprises :
      sequence of computations in a second memory bank in             accumulating, in the memory, outputs of computational
      the accelerator memory.                                   15       elements executed by the GPU in performing the first
   21. A method of performing a sequence of computations                 computation in the sequence of computations.
representing an artificial neural network, the method com-            30. The method of claim 21 , further comprising:
prising :                                                            storing, in the memory, an output of a previous compu
   receiving, at a central processing unit ( CPU ), first input          tation in the sequence of computations; and
      data acquired from an external system in real time; 20 accessing, by the GPU , the output of the previous com
   initializing, by a controller operably coupled to a graph-           putation during performance of the computation in the
      ics processing unit (GPU ), textures and shaders in a              sequence of computations.
      memory opera coupled to the GPU ;                               31. The method of claim 21 , wherein performing the first
   transferring the first input data received by the CPU to the computation comprises executing a plurality of computa
     memory operably coupled to the GPU :                       25 tional elements representing a layer of neurons in an arti
   performing, by the graphics processing unit (GPU ), a first ficial neural network.
      computation in the sequence of computations on the              32. The method of claim 31 , wherein all neurons in the
     first input data based on the textures and shaders to layer          of neurons are described by the same equation .
      generate first output data , computations in the 30 33.              The method of claim 21 , further comprising :
                                                                      acquiring the second input data with at least one of a
      sequence of computations representing respective lay
      ers of neurons in the artificial neural network , an               video camera , a microphone, or a cell recording elec
                                                                         trode .
      output of the first computation in the sequence of              34. The method of claim 21 , further comprising :
      computations representing an output of a first neuron in        loading the second input data from disk .
      a first layer in the artificial neural network ;          35    35. A system for performing a sequence of computations,
   storing , in the memory operably coupled to the GPU , the the system          comprising:
     first input data and the first output data ; and                a camera to generate input data in real time ;
   transferring second input data acquired from the external         a first memory partition ;
      system in real time into the memory operably coupled            a second memory partition operably coupled to the first
      to the GPU after the GPU starts the first computation 40           memory partition ; and
      and before the GPU starts a second computation of the           a processing unit , operably coupled to the camera , the
      sequence of computations, an output of the second                 first memory partition , and the second memory parti
      computation in the sequence of computations repre                  tion , to perform the sequence of computations on a first
      senting an output of a second neuron in a second layer            portion of the input data so as to generate a first
      in the artificial neural network .                        45      portion of output data , intermediate computations in
   22. The method of claim 21 , wherein transferring the                 the sequence of computations yielding intermediate
second input data comprises transferring the second input                results, the first portion of the output data representing
data via a bus operably coupled to the CPU .                            an   output of an artificial neural network,
   23. The method of claim 21 , further comprising:                   wherein the first memory partition is configured to trans
   transferring the first output data from the memory to 50             fer a second portion of the input data to the second
      another memory during the second computation in the               memory partition in parallel with performance the
      sequence of computations.                                          sequence of computations by the processing unit ,
   24. The method of claim 23, further comprising:                    wherein     the second memory partition is configured to
   storing intermediate results of the sequence of computa 55            transfer   a second portion of the output data to the first
      tions in the memory, and                                           memory     partition in parallel with performance the
   wherein transferring the first output data from the                   sequence     of computations by the processing unit, and
                                                                      wherein the sequence of computations represents the
      memory to the other memory occurs without transfer                 artificial neural network , each neuron in the artificial
      ring the intermediate results of the sequence of com               neural    network has an output variable assigned to a
      putations.                                                60      first texture and a second texture in the memory , the first
  25. The method of claim 23 , wherein transferring the                  texture holds a first value of the output variable com
second input data and transferring the first output data                 puted during a previous time step of the sequence of
occurs in parallel .                                                     computations and accessible to other neurons in the
  26. The method of claim 21 ,further comprising:                        neural network during a current time step of the
  storing , in a first memory partition of the memory, param- 65         sequence of computations and the second texture accu
     eters common to all of the computations in the                      mulates a second value of the output variable computed
     sequence of computations.                                           during the current time step .


            Case 7:26-mc-00318-LS                Document 6-4              Filed 08/18/26              Page 21 of 22


                                                    US RE48,438 E
                             19                                                                 20
   36. The system of claim 35, wherein the first memory                   in the sequence of computations representing layers
partition and the second memory partition are logical par                 of the neural network and yielding intermediate
titions .                                                                 results ; and
  37. The system of claim 35, wherein the processing unit is            accelerator memory, operably coupled to the at least
comprises a graphics processing unit ( GPU ) .               5
                                                                          one processing unit, to store the results of the
   38. The system of claim 35, wherein the processing unit is             sequence of computations; and
configured to receive the input data at a first rate and to             a controller, operably coupled to the at least one
perform the sequence of computations at a second rate is                  processing unit and the accelerator memory, to con
different than the first rate.                                            trol transfer of the at least a portion of the input data
   39. The system of claim 35, wherein the second memory 10            into the accelerator memory during performance of
partition is configured to transfer the second portion of the          the intermediate computations in the sequence of
output data to the first memory partition without transfer             computations by the at least one processing unit, to
ring any of the intermediate results to the first memory               control transfer at least a portion of the output data
partition .
   40. A system for executing an artificial neural network, 15         from the accelerator memory to the main memory
the system comprising:                                                 during performance of the intermediate computa
   a central processing unit (CPU ) to provide first input             tions in the sequence of computations by the at least
     data ;                                                            one processing unit, and to control performance of
  a memory, operably coupled to the CPU , to store the first           the sequence of computations by the at least one
     input data in a first partition, referenced by a first 20         processing unit.
     pointer, before computing a first layer of neurons of the    45. The computer system of claim 44, wherein the central
     artificial neural network ;                               processing unit is configured to receive the input data in
   a processing unit, operably coupled to the memory, to response to a user interaction .
     perform , during computation of the first layer of neu-       46. The computer system of claim 44, wherein :
      rons, at least one calculation on the first input data so 25 the central processing unit is configured to receive the
     as to generate first output data , the first output data         input data at a first rate ; and
     representing an output of at least one neuron in the first    the at least one processing unit is configured to perform
     layer of neurons ; and                                           the sequence of computations at a second rate different
   a controller, operably coupled to the processing unit and          than the first rate.
      the memory, to :                                          30 47. The computer system of claim 44, wherein the main
     store the first output data in a second partition of the memory is configured to store a copy of the output data
        memory, the second partition referenced by a second stored in the accelerator memory.
        pointer, and to swap the firstpointer with the second      48. The computer system of claim wherein an output
       pointer at the end of the computation of the first layer of at least one computation in the sequence of computations
        of neurons, such that the firstoutput data becomes an 35 represents an output of at least one neuron in an artificial
        input for a second layer of neurons of the artificial neural network .
        neural network,                                              49. The computer system of claim 44, wherein accelerator
     transfer the first output data to another memory during memory comprises:
        computation of the second layer of neurons, and              a first memory partition to store parameters common to
     dictate an order of execution of instructions to the 40            all of the computations in the sequence of computa
       processing unit to perform the computation of the                tions ; and
       first layer of neurons .                                      a second memory partition to store data specific to at least
  41. The system of claim 40 , wherein the processing unit              one computation in the sequence of computations.
comprises a graphics processing unit .                               50. The computer system of claim 44, wherein the con
  42. The system of claim 40 , wherein the controller is 45 troller is configured to transfer the output data from the
configured to send instructions for performing the at least accelerator memory to the main memory without transfer
one calculation to the processing unit.                           ring any of the intermediate results from the accelerator
  43. The system of claim 40 , wherein the memory further memory to the main memory so as to reduce data transfer
comprises :                                                       via the bus.
  a third partition to store internal variables; and           50    51. The computer system of claim 44, wherein the con
  a fourth partition to store data used as input at a troller is configured to transfer at least a portion of the
    particular layer of neurons of the artificial neural output data from the accelerator memory to the main
     network.                                                     memory after the at least one processing unit has begun to
  44. A computer system , comprising:                             perform another sequence of computations.
  a central processing unit to receive input data acquired 55 52. The computer system of claim 51 , wherein the con
    from an external system ;                                     troller is configured to initiate transfer of the at least a
  main memory, operably coupled to the central processing portion of the input data and to transfer the at least a portion
     unit via a bus, to store the input data received by the of the output data in parallel with performance of at least
     central processing unit;                                     one computation in the other sequence of computations by
  an accelerator, operably coupled to the central processing 60 the at least one processing unit .
     unit and the main memory via the bus, to receive at             53. The computer system of claim 44, wherein the con
     least a portion of the input data from the main memory , troller is configured to control execution of the sequence of
     the accelerator comprising :                                 computations by the at least one processing unit .
     at least one processing unit to perform a sequence of           54. The computer system of claim 44, further comprising :
        computations representing an artificial neural net- 65 at least one of a video camera , a microphone, or a cell
        work on the at least a portion of the input data so as          recording electrode, operably coupled to the central
        to generate output data, intermediate computations              processing unit, to acquire the input data in real time .


       Case 7:26-mc-00318-LS                   Document 6-4       Filed 08/18/26    Page 22 of 22


                                                  US RE48,438 E
                            21                                                 22
   55. The computer system of claim 1 , wherein the control
ler is configured to inform the central processing unit that
the sequence of computations is finished .
   56. The computer system of claim 1 , wherein the control
ler is configured to reduce a processing load on the central 5
processing unit.
   57. The computer system of claim 1 , wherein the control
ler is configured to reduce interactions between the central
processing unit and the accelerator.
                                                            10