Welcome to WsxMall | Register
BOM
WsxMall > Industry Information > Intelligent Driving Chip Top20 Ranking

Intelligent Driving Chip Top20 Ranking

WsxMall 2023-12-28 17:11:21 130 Related Key Words: Smart driving chip

The ranking of smart driving chips is not easy to look at only AI computing power. The CPU, storage bandwidth, power consumption and AI computing power value are as important. This will be analyzed in detail.The CPU computing power is also very important. Smart driving system software is extremely complicated and consumes a lot of CPU computing resources. The software system includes many middleware such as SOME/IP, adaptive AutoSAR, DDS, ROS, etc. The basic software includes the custom Linux & Nbsp; BSP, OS abstraction layers, virtual machines, as well as memory management associated with the bottom hardware, various drivers, various communication protocols, and so on.In addition, the path planning, high -precision maps, and behavioral decision -making in the application layer also consume a lot of CPU resources. At the same timeTask.The CPU is an absolute core. AI is a subsidiary of the CPU. It is only used in image feature extraction, classification, BEV transformation, vector map mapping or spatial distribution.

The weight of the ranking is AI computing power, storage bandwidth, CPU computing power, GPU computing power, manufacturing technology.The storage bandwidth and AI computing power are the same weight. The GPU is also icing on the cake. Most of the car AI processing parts can only correspond to INT8 -bit data, and the GPU can correspond to FP32 data, and sometimes it may have a great effect.The actual AI computing power figure is completely a black box, the operation space is extremely great, and the reference is not significant.The most accurate measurement power is the number of MAC arrays. Google's TPU & Nbsp; V1 is 65,000 FP16 MACs, and the operating frequency is 0.7GHz. Then the computing power is 65000*0.7g*2 = 91TOps.Tesla's first -generation FSD NPU, each NPU is 9216 INT8 MAC, the operating frequency is 2GHz, and the computing power is 2*2*2G*9216 = 73.7TOPS.In terms of manufacturing technology, the more advanced, the lower the power consumption.

Intelligent driving chip top20

Image source: Public information sorting

How to calculate the storage bandwidth, the chip itself has a storage manager. This is usually part of the CPU. There are two points that determine the storage bandwidth.Bandwidth, the highest storage bandwidth of LPDDR is generally 256 bits, GDDR can reach 384 bits, HBM can reach 4096 or even 8192 bits. These are related costs. When designing chips, manufacturers will find a balance between cost and performance.The manufacturer's heavy cost, then 64 bits or even 32 bits, with some emphasis on performance, such as the real AI chip, without exception, HBM, and the cost is more than $ 1,500.

Common car memory performance and price comparison

Image source: Public information sorting

The above table shows the comparison of common car memory performance and price. Obviously, a price is cost.Nvidia H100 is the largest buyer of HBM3, and the purchase price of each GB is about $ 14.There is also a need to point out that there are currently no car -level GDDR6 storage chips.

In addition to Baidu and Tesla, smart driving chips use LPDDR.

Parameters of LPDDRs in the past

Image source: Public information sorting

Storage bandwidth equal to the storage position of the CPU. The wide multiplication of the Data Transfer Rate of the memory, DDR (MT/S) and then divide the GB with 8 converted by 8./S, then the storage bandwidth is 256*6400m/8 = 204.8GB/s.8 = 51.2GB/s.

The important reason for the storage bandwidth is the ROOF-LINE model. The ROF-LINE Model is solved. It is the upper limit of the theoretical performance of the calculation platform with the calculation of the calculation amount and the band-width of D.How much is E.

The calculation amount of the model refers to the input of a single sample (for CNN, which is a image). The model is performed in a complete front -directional propagation, which is the time complexity of the model. The unit is FLOPS.Visit stock: refers to the input of a single sample. The model completes the total amount of memory exchange occurred during the propagation process, that is, the space complexity of the model.Ideally (that is, without considering the cache on the film), the number of models of the model is the sum of the memory occupation (Kernel Mem) of the weight parameters of each layer of the model and the memory occupation (Output Mem) output from each layer.In addition to the amount of calculation, the calculation strength I (INTENSITY) can be obtained in except for the exist. It indicates that during the calculation process, how many floating -point operations are used for each Byte memory exchange.The unit is FLOP/BYTE.The number of floating -point operations (theoretical values) per second that the model can reach on the computing platform.The unit is FLOP/S, which is P.

The computing power determines the height of the roof (green line segment), and the bandwidth determines the slope of the eaves (red line segment)

The theoretical performance of the model calculation cannot naturally exceed the maximum theoretical performance of its hardware. If there is an abnormal consumption power model, the computing power required exceeds the theoretical performance of the calculation platform, then the utilization rate of the computing platform is 100%, and it is also the utilization rate, and it is also 100%, and it is also 100%, and it is also 100%.It is the red line segment. At this time, the risk is to process the frame rate of the image or the FPS will not reach the target frame rate. For smart driving, the mainstream frame rate is 30fps. Low -speed intelligent driving can be reduced by a little bit.higher.Because the required computing power is too high, the computing platform's full load operation cannot be adapted, and the frame rate will decline. At this time, there will be risks if driving at high speed. Generally speaking, manufacturers will not recommend the model of computing power far exceeding the model of the theoretical performance limit.Essence

In the green line segment below 100%utilization, the size of model theoretical performance P is completely determined by the upper limit of the bandwidth (slope of the eaves) and the calculation intensity of the model itself.Memory-Bound state.It can be seen that under the premise that the model is in the bandwidth bottleneck interval, the bandwidth of the computing platform is that the steeper the eaves, or the larger the calculation strength I of the model, the larger the theoretical P.The lower the slope, it means that even if the calculation intensity increases rapidly, the increase in the computing platform's computing power is still very slow, and the utilization rate of the calculation platform is very low.The utilization rate may also be less than 50%. In other words, the storage bandwidth determines the performance utilization rate of the calculation platform. Therefore, the importance of storage bandwidth is no less than computing power, or even higher than computing power.This is also the main reason why Tesla's second -generation FSD ranks second. The bandwidth of GDDR6 has an overwhelming advantage over LPDDR.

Tesla's second -generation FSD

Tesla's second -generation FSD adopts Samsung's 7 -nanometer. The reason why Samsung founded Samsung is mainly prices and geographical factors.The factory is low -efficiency. Construction started in 2020 and is expected to be put into production until 2025. Samsung's Texas Austin's second -generation factories have been completed and put into operation in just two years, and Tesla headquarters is very close to Austen.The first generation of FSD used Samsung's 14 nanometer technology. Wikichip data showed that the transistor density of Samsung 7nm LPP HD high -density Cell scheme is 95.08 mtr/mm², while the transistor density of HP high -performance scheme is 77.01 mtr/mm²; Samsung 14The crystal density of the nano UHP solution is 26.22 mtr/mm², and the HP scheme crystal density is 32.94 mtr/mm². Basically, Samsung 7 nanometers are more than three times that of 14 nanometer density, which means that Tesla can be stuffed at least more than 3 times more than 3 times more than 3 times more than 3 times more than 3 times more than 3 times more than 3 times more than 3 times.For the Mac array, the performance of AI can be three times. The AI performance of the first generation of FSD is 73.7TOPS@int8, and 3 times is 221.1.The area of the FSD chip is obviously larger than the first generation, and the NPU increases to 3, so the estimated calculation power is around 500TOPS.Tesla's second -generation FSD has also greatly strengthened the CPU, using Samsung Exynos 20 core configuration, which also shows that the CPU is important in smart driving.

There are not many people who are familiar with Anba's CV3. Its storage bandwidth supports the highest LPDDR5X, and it is the highest 256 bit. It is manufactured by Samsung's 5 -nanometer process. It is currently supported by the German mainland car company.

Anba CV3-AD internal frame map

Image source: Ambarella

Anta CV3-AD includes the 16-core Coretex-A78Ae, and the CPU computing power is also extremely high.ASIL-B certification has also been passed.AI computing power is equivalent to 500TOps.Nvidia's position is 256 bits, Tesla and Mobileye are mostly 128 bits. The journey 6 has not been announced.

Baidu's Kunlunxin 2 rarely knows that in fact, this cannot be counted as Baidu. It is the product of the independent part of the Baidu chip. The full name of the company is Kunlunxin (Beijing) Technology Co., Ltd.Independent financing was completed in April of the year, and the first round of valuation was about 13 billion yuan.On November 29, 2022, on the opening day of Baidu Apollo Day Technology, the second -generation Kunlun core has been fully adapted on the driving system of Baidu's driverless vehicle Robotaxi, which runs normally in the high -end autonomous driving system.In 2011, Kunlunxin Technology was officially independent and began to engage in AI computing related work. Early, the FPGA chip was used to accelerate the calculation of AI.From 2011 to 2015, Kunlunxin Technology deployed more than 5,000 FPGA chips for Baidu Data Center. By 2017, it deployed more than 12,000 FPGA chips.In 2018, he decided to develop the AI chip and officially launched the research and development and design of the Kunlun core series of products.In 2020, the first generation of Kunlunxin began to deploy large -scale. In 2022, the second -generation Kunlunxin was deployed and landed on large -scale in the fields of data centers, industrial fields, and autonomous driving.The first generation of Kunlunxin is a 14 -nanometer artificial intelligence chip. This chip uses advanced HBM memory and 2.5D packaging. The chip just mass production has deployed more than 20,000 pieces in Baidu Data Center.One year later, the second -generation Kunlun core was mass -produced. It adopted a more advanced 7 -nanometer process and the second -generation structure of XPU. It is also the first AI chip in the industry to use GDDR6 high -speed memory technology.Kunlunxin Technology is developing a more advanced third -generation AI chip. For high -end autonomous driving systems, in the future, we will consider launching customized car regulations and high -performance SOC (system -level chips).

Nvidia pays more attention to the first direction of the storage system, and the entire line is the highest 256 bit.The SA8650 of Qualcomm is very similar to the SA8255 in the cockpit field. The CPU and GPU are basically the same. The AI computing power has been specially strengthened. The storage position is a relatively rare 96 bit. The SA8650 is replaced by the previous generation SA8540P.The section adds 4 Cortex-R52 kernel.Mobileye attaches great importance to costs and never announces its storage bandwidth and support storage types. It can only guess.Although Xavier is an early product, the storage position is the highest 256 -bit, so the ranking is very high.

Demonstration: The views and data of this article are for reference only, and there may be deviation from the actual situation.This article does not constitute investment suggestions. All views and data in the article only represent the author's position, and do not have any guidance, investment and decision -making opinions.

Hot -selling model

Product Index :