E8 Back-End Physical Design and Tapeout Preparation
You have already integrated the NPC into the SoC and verified the functionality of the SoC system. You have also measured the performance of the NPC on the SoC. In addition, you have completed the synthesis of the NPC and generated the netlist. You are now very close to the goal of tape-out. The remaining work is to carry out the back-end physical design and generate a tapeout-ready GDS layout from the netlist.
It is important to note that, for the chip to be taped out, we will provide a correct version of ysyxSoC. The layout you generate will be integrated into the ysyxSoC we provide, together with the layouts generated by other students. In other words, all students will share the same SoC in the chip that will eventually be taped out. Therefore, you only need to perform synthesis and back-end physical design for the NPC itself. You do not need to include ysyxSoC in this process. Your previous integration of the NPC into ysyxSoC was primarily intended to verify that the NPC could be correctly simulated within the ysyxSoC environment.
However, before proceeding with the back-end physical design, you still need to complete some front-end preparation work to identify and eliminate potential issues in the NPC.
Front-End Preparation Work
Open Up the NPC Address Space
The SoC used in the tapeout environment will integrate more peripherals than ysyxSoC, allowing your NPC to provide richer demonstrations after the chip is fabricated and returned. However, if the NPC intercepts accesses to the corresponding address spaces in advance, it will not be able to access these peripherals after tapeout. Therefore, we recommend that you open up the entire address space in the NPC and allow read and write requests to all addresses to be sent through the NPC's external SimpleBus interface.
Open Up the NPC Address Space
Check your NPC implementation to make sure that the entire address space is open. If your NPC intercepts accesses to any part of the address space, you need to modify the NPC code accordingly.
Remove Falling-Edge-Triggered Clocks
Using both rising-edge- and falling-edge-triggered clocks makes timing closure more difficult and increases the complexity of back-end physical implementation. In general, there is no need to use falling-edge-triggered clocks inside a processor.
Remove Falling-Edge-Triggered Clocks
Check your code to see whether it contains any falling-edge-triggered clocks. If there are any, you need to modify the corresponding modules to remove them.
Remove Latches
In synchronous sequential logic circuits, access to all storage elements is controlled by the clock.
However, latches are not driven by clock edges, making them difficult for timing analysis tools to analyze. Therefore, they are generally not used in synchronous sequential logic circuits.
Remove Latches
In ICsprout55, latch cells are named with the prefix LAT*. You need to check whether the synthesized netlist contains any latch cells. If so, you can locate the corresponding RTL code using the source file locations provided in the comments of the netlist, and remove the latches from the design.
Chisel Bonus
Verilog code generated by Chisel describes synchronous sequential logic circuits, so if you use Chisel for development, you don't need to worry about generating latches.
Naming Modifications
The New Tapeout Integration Scheme Does Not Require Naming Modifications
Under the new tapeout integration scheme, after completing the physical design, everyone will integrate their layouts into a single SoC. In this layout-based integration scheme, the names used in your RTL code will not conflict with those of other students. Therefore, in principle, you do not need to modify the names in your RTL code, and you do not need to complete this section.
However, we are gradually transitioning to the new tapeout integration scheme. During this transition period, this section of the guide will be retained for reference when needed.
Under the old integration scheme, we would integrate code from multiple students into the same SoC. If multiple designs used the same module names, the EDA tools would report module redefinition errors. To resolve this issue, we required the following modifications:
Merge the CPU code into a single
.vfile namedysyx_8-digit-student-ID.v, such asysyx_22040228.v.- On Linux, this can be done using the
catcommand:
$> cat CPU.v ALU.v regs.v ... > ysyx_22040228.v- On Linux, this can be done using the
Rename the CPU top-level module to
ysyx_8-digit-student-ID, such asysyx_22040228.Add the prefix
ysyx_8-digit-student-ID_to all variables defined using`definewithin the CPU. For example,`define SIZE 5should be changed to`define ysyx_22040228_SIZE 5. If you use Chisel, you do not need to make these modifications.Add the prefix
ysyx_8-digit-student-ID_to all module names within the CPU. For example,module ALUshould be changed tomodule ysyx_22040228_ALU.
Chisel Bonus (2)
If you use Chisel, you can use Chisel's module prefixing feature to automatically add module name prefixes. For details on how to use it, refer to the relevant documentation and examples. If you use Verilog/SystemVerilog, automatic module name prefixing is currently not supported, so you will need to add the prefixes manually.
Static Code Analysis
Verilator can perform static analysis on Verilog code and identify potential issues that may lead to errors. Fixing these issues can help improve the correctness and robustness of your code. Specifically, you can use Verilator's --lint-only option to perform the checks described above.
To have Verilator perform as many checks as possible, you can also add the -Wall option. The Verilator documentation provides the meaning of all warnings. In principle, you should resolve as many warnings as possible to make your code more robust and规范. The following two types of warnings are exceptions:
- The
DECLFILENAMEwarning is unrelated to the logic of your code, so you can suppress it with-Wno-DECLFILENAME. - For warnings related to
UNUSED, some are caused by unused input ports and can be ignored once you have confirmed that they are intentional. Others may indicate that your code is not written properly, so you should carefully investigate the cause. If you decide to suppress a warning, you should understand what problems it might cause and take responsibility for your decision.
Use Verilator to Check the RTL Code
Use Verilator's --lint-only option to check your code. Review the information reported by Verilator and determine, one by one, whether the corresponding RTL code needs to be modified.
Chisel Bonus (3)
Verilog code generated by Chisel is generally compliant with coding conventions, but it may still produce UNUSED warnings related to signals and bit widths. If you use Chisel for development, this part of the work should be somewhat easier for you. Nevertheless, we still recommend carefully checking every warning.
Flip-Flop Reset and Four-State Simulation
The Necessity of Resetting Flip-Flops
System startup can be divided into cold start and warm reset. A cold start refers to powering on and resetting the system from a powered-off state, while a warm reset refers to resetting the system while it is already powered on. As we know, the state of a circuit is determined by the states of its sequential logic elements. Therefore, we expect the circuit to enter a known, correct state after reset and then begin operation from that state. To achieve this, reset values must be specified for the sequential logic elements.
One approach is to reset all sequential logic elements, ensuring that the system enters exactly the same state after reset, regardless of whether it is a cold start or a warm reset. However, for circuits containing random-access memory, this is difficult to achieve because RAM generally does not provide a reset function, whether it is SRAM integrated inside the CPU or DRAM external to the CPU. This means that during a warm reset, the state of the RAM's memory array is not affected by the reset signal, and the data already stored in memory is carried over into the next reset. Therefore, the computer's hardware and software systems must ensure that the system can operate correctly regardless of what data is stored in RAM after a reset. In other words, the system's behavior during reset must be designed to be independent of the contents of the RAM.
If we exclude random-access memory, the remaining sequential logic elements are essentially flip-flops. Resetting all flip-flops can certainly increase the likelihood that the system will enter a correct state, but it may also introduce additional delay and some area overhead. On the one hand, the reset signal needs to be routed to all flip-flops during physical design, resulting in greater propagation delay for the reset signal. On the other hand, the reset functionality of each flip-flop requires additional logic gates, which also consume a certain amount of area.
In real-world projects, as long as the system is guaranteed to operate correctly after reset, only the minimum necessary subset of flip-flops is typically reset. For flip-flops that do not need to be reset, their state is unaffected by the reset signal, so they retain their previous state when the system is reset. As long as the value stored in a flip-flop does not affect the system's behavior after reset, that flip-flop does not need to be reset.
Typically, there are two main categories of flip-flops that do not need to be reset:
Registers that can be initialized by software before being used. This means that these registers must be software-visible; in other words, they are defined by the ISA specification. For example, for general-purpose registers other than the zero register, the RISC-V specification states that their values after reset are undefined. Therefore, software should write to these registers before reading them, ensuring that they contain meaningful values. In fact, specifying the initial values of these registers as undefined in the ISA essentially moves their initialization from the hardware layer to the software layer, thereby reducing circuit area and delay.
Flip-flops in the datapath associated with control signals. Typically, control signals such as
validandenare used to indicate whether the corresponding data signals are valid. When the data is invalid, downstream modules connected to these control signals generally neither use nor store the corresponding data. Therefore, most flip-flops in the datapath do not need to be reset, as long as the circuit guarantees that the corresponding flip-flops contain valid data by the time their associated control signals indicate that the data is valid.
Resetting flip-flops that do not need to be reset at most wastes some area and increases the propagation delay of the reset signal; it does not affect the correctness of the circuit. However, failing to reset flip-flops that do require reset may cause the circuit to malfunction after reset. Therefore, the latter situation must be avoided.
Four-State Simulation with Icarus Verilog
Previously, we have been using Verilator for simulation. Verilator is a two-state simulation tool, where all signals can only have two possible values: 0 and 1. As a result, every flip-flop always has a definite value during reset, making it difficult to detect the issues described above.
In an RTL simulator that supports four-state simulation, the unknown state X can be used to represent the initial value of a flip-flop that has not been reset. When an X signal participates in logic-gate operations, the result may also be X, allowing the unknown state to propagate through the circuit. If a flip-flop that should be reset is not reset, the X signal may propagate extensively through the circuit, causing the values of other flip-flops to also become X. These signals can no longer be represented by precise values, and eventually the circuit may converge to a state different from the expected correct state. As a result, programs running on the processor may fail to produce the expected results. Conversely, if there are no flip-flops that require reset but are left unreset, the propagation of X signals will be contained by control signals. As valid data is gradually written into the circuit, the X signals will gradually disappear, and the circuit will eventually converge to the expected correct state.
Icarus Verilog, also known as Icarus Verilog, is an RTL simulator that supports four-state simulation. We can use it to check whether there are any flip-flops that should be reset but are not properly reset.
You already obtained Icarus Verilog when you downloaded the open-source CAD tool suite, oss-cad-suite.
Try Using Icarus Verilog
Refer to the Icarus Verilog documentation and run a simple example using Icarus Verilog.
Since we will ensure the correctness of the ysyxSoC used for tapeout, you only need to simulate the NPC independently with Icarus Verilog for now, and run the minirv-npc AM program on it.
However, since Icarus Verilog does not currently support DPI-C, the way the simulation is driven is also different from Verilator. Therefore, we need to make some modifications to the RTL code and simulation environment. When compiling with Icarus Verilog, the __ICARUS__ macro is automatically defined. You can therefore wrap these modifications in conditional compilation directives, allowing your code to support simulation with both Verilator and Icarus Verilog.
Implementing Memory Access via VPI
In Icarus Verilog, we can use the VPI mechanism to enable communication between Verilog and C code. You can first read the relevant documentation to learn how to use the VPI mechanism in Icarus Verilog.
After gaining a basic understanding of the VPI mechanism, we can consider how to implement memory access through VPI in Icarus Verilog. One approach is to use VPI to register two system tasks named $pmem_read and $pmem_write. We can then replace calls to the DPI-C functions pmem_read() and pmem_write() in the Verilog code with calls to the system tasks $pmem_read and $pmem_write. The handler functions (written in C) registered for these system tasks can then call the previously implemented C functions pmem_read() and pmem_write(). These C functions will ultimately access the array used to simulate memory in the C code.
However, there are still some issues that need to be addressed:
- How can we retrieve the arguments passed from the Verilog code in the system task's handler function?
- How can we make
$pmem_readreturn the value read bypmem_read()? - How can we load the program to be executed into the array before starting the simulation?
The introduction to the VPI mechanism in the Icarus Verilog documentation does not cover these topics. You will need to STFW to learn more about the capabilities of VPI. If necessary, you can also consult the Verilog standard manual for the information you need.
Check the Impact of Unreset Flip-Flops on Circuit Behavior
In addition to implementing memory access through the VPI mechanism, you also need to make the following modifications:
Make the NPC start executing the program from
0x80000000. You can consider, but are not limited to, the following approaches:- Change the reset value of the PC to
0x80000000. - Keep the PC reset value at
0x30000000, but place a few instructions at0x30000000that immediately jump to0x80000000, allowing the NPC to execute the AM program fromriscv32e-npc.
- Change the reset value of the PC to
After detecting an
ebreakinstruction, you can use$display()to print relevant information and terminate the simulation with$finishor$fatal.Disable the DiffTest mechanism.
Write a simple simulation top module that instantiates the clock and reset signals and drives the NPC simulation. For example:
module iverilog_top(); reg clock; reg reset; initial begin clock = 0; reset = 1; # 10 reset = 0; end always # 1 clock = ~clock; Top top(clock, reset); endmoduleHere, the
Topmodule contains the NPC and memory modules, with the memory module implementing memory access through the VPI mechanism.
After making the above modifications, you can try compiling your design with Icarus Verilog. However, please note the following:
- Although Icarus Verilog can automatically identify the top-level module, we still recommend explicitly specifying the top-level module with the
-soption to avoid Icarus Verilog selecting an unintended top-level module. - Icarus Verilog supports a smaller subset of Verilog/SystemVerilog syntax than Verilator. We recommend using the newer language standard with the
-g2012option. However, you may still encounter cases where the code compiles successfully with Verilator but fails to compile with Icarus Verilog. You will need to modify your code according to the error messages.- In particular, if you use Chisel, you also need to add the
--lowering-options=disallowLocalVariables,disallowPackedArraysoption tofirtoolto disable some language features that Icarus Verilog does not support.
- In particular, if you use Chisel, you also need to add the
- You may also encounter the following error message, which can be safely ignored:
sorry: constant selects in always_* processes are not currently supported (all bits will be included). - Since DiffTest is currently unavailable, if the simulation results differ from what you expect, you will need to diagnose the problem using waveforms. You can read the Icarus Verilog documentation on viewing waveforms to learn how to generate waveforms with Icarus Verilog.
- However, we recommend that you first use Verilator and DiffTest to eliminate as many potential issues as possible, and then use Icarus Verilog for simulation.
- You can also try implementing DiffTest in the Icarus Verilog simulation environment using the VPI mechanism.
Check X Signal Propagation with Four-State Simulation
Simulate the NPC with Icarus Verilog and run programs such as microbench. If the program fails to run correctly, you need to modify the corresponding RTL code to resolve the X signal propagation issues.
Gate-Level Netlist Simulation
Verilog is essentially an event-driven hardware modeling language, and synthesizable Verilog is only a subset of the language. Therefore, a Verilog simulator may accept code that cannot be synthesized. This is not a bug in the Verilog simulator; it is simply following the behavior defined in the Verilog standard manual. Clearly, when designing RTL, we do not want to write such code.
One way to check whether a project contains non-synthesizable code is to simulate the synthesized circuit. If the project contains non-synthesizable code, the synthesized circuit may behave differently from the original RTL design, helping us identify the corresponding issues.
Synthesis converts RTL code into logic-equivalent standard cells. The file that describes the connections between these standard cells is called a netlist. In a netlist, standard cells are instantiated as modules. Therefore, to simulate a netlist, we need to provide behavioral simulation models for these standard cells. The ICsprout55 PDK project already provides the corresponding simulation models, which are located at:
icsprout55-pdk/IP/STD_cell/ics55_LLSC_H7C_V1p10C100/ics55_LLSC_H7CH/verilog/ics55_LLSC_H7CH.vicsprout55-pdk/IP/STD_cell/ics55_LLSC_H7C_V1p10C100/ics55_LLSC_H7CR/verilog/ics55_LLSC_H7CR.vicsprout55-pdk/IP/STD_cell/ics55_LLSC_H7C_V1p10C100/ics55_LLSC_H7CL/verilog/ics55_LLSC_H7CL.v
Update the PDK
We have updated the simulation models in the ICsprout55 PDK to fix an issue that prevented Verilator from simulating them correctly. If you obtained the ICsprout55 PDK before 14:30:00 on August 21, 2026, please delete the existing PDK and obtain it again. For details, refer to the instructions for obtaining the PDK earlier in this guide.
We will first consider using Verilator for gate-level netlist simulation. The specific goal is to replace the pre-synthesis NPC RTL module with the synthesized NPC netlist module:
+--------------------------------+
| Top |
| +-------------+ +-----+ |
| | NPC-netlist | <---> | Mem | |
| +-------------+ +-----+ |
+--------------------------------+
Specifically, we need to simulate the following files:
- The synthesized NPC netlist file. You need to use the netlist file named
xxx_Synthesis_sim.v.gz. Since it is compressed, the RTL simulator cannot read it directly, so you also need to use thegunzipcommand to decompress it. - The memory module that implements memory access using DPI-C.
- The behavioral simulation model files for the standard cells mentioned above.
It is worth noting that ECC generates two netlist files after synthesis, each intended for a different purpose. The file named xxx_Synthesis_sim.v.gz is the simulation netlist. Its top-level module ports are exactly the same as those of the top-level module before synthesis, allowing it to directly replace the pre-synthesis RTL module and making gate-level netlist simulation straightforward. The file named xxx_Synthesis.v.gz is the netlist for back-end physical design. Its top-level vector ports are split into multiple single-bit ports, making it easier for back-end EDA tools to process.
The behavioral simulation model files for the standard cells also contain statements related to timing checks. However, we only need to perform functional simulation for now and do not need to consider these statements. Therefore, we need to add the following compilation options to the Verilator command line:
--timescale "1ns/1ns"--no-timing-D__VERILATOR__-Dfunctional
Since the synthesis process cannot recognize DPI-C, we cannot use the DPI-C mechanism to interact with C code from within the netlist. In addition, since the general-purpose register file has been synthesized into individual flip-flops, it is difficult to reconstruct the state of the register file by combining these flip-flops in the netlist. Therefore, it is difficult to use DiffTest to locate problems during netlist simulation. These two issues introduce new challenges for debugging the netlist simulation. For this reason, we recommend that you first perform RTL simulation with Verilator and use DiffTest to eliminate as many potential issues as possible before proceeding with netlist simulation.
Perform Netlist Simulation with Verilator
Follow the requirements above to perform netlist simulation with Verilator, and try running programs such as microbench on the NPC netlist.
Chisel Bonus (4)
Verilog code generated by Chisel is synthesizable. If you use Chisel for development, netlist simulation will most likely succeed directly. Nevertheless, we still recommend that you do not skip netlist simulation for this reason.
Finally, we also need to perform netlist simulation with Icarus Verilog to check whether the synthesized netlist contains any X-state signal propagation that prevents the simulation from completing successfully. Since the netlist consists only of standard-cell instantiations, and you have already implemented memory access through the VPI mechanism, no further code modifications are required to simulate the netlist with Icarus Verilog.
Simulate Netlist with Icarus Verilog
Simulate netlist with Icarus Verilog, and try running programs such as microbench on the NPC netlist.
Check Your Code Again
If you modified your code during the process above, repeat the checks described earlier to make sure that you did not introduce any code that violates the requirements while making these modifications.
Back-End Physical Design
EDA Tools May Be Updated Frequently
ECC and ECOS Studio were first released publicly in August 2026. Since their release, the ECOS team has received a large amount of feedback and suggestions for improvement. Therefore, ECC and ECOS Studio are expected to undergo frequent updates in the second half of 2026, both to fix various issues and to support more useful and convenient features. The relevant sections of this guide will be updated accordingly. We hope to provide everyone with a better experience as soon as possible. If you encounter any bugs while using ECC or ECOS Studio, or if you have any suggestions for either tool, feel free to submit your feedback through GitHub Issues:
After initially verifying the functional correctness of the netlist, you can proceed with the back-end physical design.
Getting ECOS Studio
We use the ECOS Studio GUI tool to perform the back-end physical design. Back-end physical design is closely related to the physical aspects of a chip. If you are new to this area, you can use the visualization features in ECOS Studio to gain a better understanding of the physical design process.
First, obtain ECOS Studio v0.1.0-alpha.9. After opening the link, find the Assets section at the bottom of the page and click the link for the file with the .AppImage extension to download it. An .AppImage is an executable file. After downloading it, you first need to grant it executable permission:
chmod a+x ECOS-Studio_..._x86_64.AppImage
Replace ... in the filename above with the appropriate version number. After granting executable permission, you can run it directly:
./ECOS-Studio_..._x86_64.AppImage
Performing Back-End Physical Design with ECOS Studio
After opening ECOS Studio, click Backend Design, then click New Workspace. In the wizard that appears, configure the following options:
Project Setup- Click
Create Project- Enter
Project Name - Specify
Project Parent Path
- Enter
- Click
Continue
- Click
Basic Info- Enter
Workspace Name - Click
Continue
- Enter
Flow Setup- Select
Floorplanfrom theSTART STEPdropdown menu- This skips the
Synthesisstep because we have already performed synthesis with ECC and obtained the netlist.
- This skips the
- Click
Continue
- Select
Design Files- Click
Import Verilogand select the netlist file generated by ECC, such aslight/runs/default/Synthesis_yosys/output/light_Synthesis.v.gz.- Note that you should select the netlist intended for back-end physical design, i.e. the file whose name does not contain
_sim.
- Note that you should select the netlist intended for back-end physical design, i.e. the file whose name does not contain
- Click
Continue
- Click
PDK Config- Check the
Process Design Kitinformation box.- If ECOS Studio automatically detects the path to the ICsprout55 PDK, verify that the path is correct.
- If ECOS Studio does not detect it automatically, click
Import PDKand manually select the path to the ICsprout55 PDK.
- Select
Default Configin theConfig Modebox. - Click
Continue
- Check the
Spec Setting- Enter
Design Name. This can be different from the top-level module name in your code. - Enter
Top Module Name. This must match the top-level module name in your code. - Enter
Clock Signal Name. This must match the name of the clock signal in the top-level module. - The other parameters can be left at their default values.
- However, you may want to pay attention to
Frequency max [MHz]andOrigin Core Utilization.- The former specifies the target clock frequency of the chip.
- The latter specifies the target core utilization, i.e. the percentage of the core area occupied by the design.
- However, you may want to pay attention to
- Click
Create Workspace.
- Enter
After completing the above steps, ECOS Studio will create the workspace and open the Step Configuration window. For now, we do not need to configure the details of each individual step. Simply click the close button in the upper-right corner. After closing the Step Configuration window, you will be on the DASHBOARD page of the workspace. There are many other subpages available on the left side of the interface, but we will not explore them for now. In the upper-right Flow status section of the DASHBOARD page, there is a Start button. Click it to run all the steps sequentially.
After all steps have completed successfully, you can inspect the results of the back-end physical design. The Key Metrics section in the lower-left corner of the DASHBOARD page displays several key metrics of the design. Some of the key metrics include:
Die Area— The chip area. A lower utilization target results in a largerDie Area.Frequency— The operating frequency achieved after completing the back-end physical design.- This value will usually be significantly lower than the frequency reported during synthesis, mainly for two reasons:
- The frequency reported during synthesis is calculated based on the delays of the standard cells themselves, whereas the frequency after back-end physical design also takes into account the delays of the interconnects between standard cells.
- The frequency reported during synthesis is calculated assuming that the chip operates under typical conditions, whereas the frequency after back-end physical design is calculated assuming that the chip operates under worst-case conditions. Therefore, the post-layout frequency is considerably more pessimistic than the synthesis result.
- This value will usually be significantly lower than the frequency reported during synthesis, mainly for two reasons:
After all steps have completed successfully, you can export the signoff package. Specifically, click File → Export Signoff Package in the menu bar at the top-left corner. In the Signoff Package Review window that appears, review the information about the resources to be exported. The Risk Details section may report that config.macro_locations is missing. This is expected because our current design does not contain any macro cells. Other than this, the Risk Details section should not report any other warnings or errors. Click the Export Package button in the lower-right corner and select an export path to complete the export of the signoff package.
If you cannot click Export Signoff Package in the menu bar, or cannot click the Export Package button in the Signoff Package Review window, it is usually because the results of the back-end physical design have not met the required sign-off criteria. In this case, locate the Checklist section of the workspace, which is to the left of the Flow status section. Click Sign-off details in the lower-right corner of the section to open a new window and check which items in the checklist have not met the requirements. Then, based on the results, modify some of the parameters in the workspace's Spec Setting:
- If the STA-related checks fail, try setting a lower target frequency.
- If the DRC- or LVS-related checks fail, try setting a lower utilization target.
Specifically, you can click File → New Workspace in the menu bar to create a new workspace. Alternatively, you can click File → Update Workspace to update the parameters of the current workspace. You can then rerun the back-end physical design in the new or updated workspace until the design results meet the required criteria.
Complete the Back-End Physical Design of the NPC with ECOS Studio
Follow the procedure above and try to perform the back-end physical design for the NPC.
Explore ECOS Studio
As an optional exercise, try to find the highest utilization target at which a signoff package can still be successfully generated.
There are many new concepts in back-end physical design, which we will not cover in detail here. We will explain these concepts in greater detail during the D stage. If you would like to learn more about ECOS Studio, you can refer to the ECOS Studio User Guide.
