The last post was mostly tooling related; we didn't make progress towards our goal, but we set ourselves up for success with a good debugging setup. Now we will get back on track and explore the features of the stm32l432kc that will be needed, and figure out how to use them.

I want to preface this post with a disclaimer; I don't know what I'm doing. I am learning all of this stuff as I go along. You should not take any of my struggles as a model or a blueprint for your own projects. And if you see me making a stupid mistake, please, let me know!.

Initializing the data and bss segments

While my program is persisted in flash memory, the data segment, which can contain global and static variables, needs to be in SRAM so data can be modified without needing to reprogram a whole page of flash. So I added an _init routine which initializes the data and bss segments using symbols about their locations that are placed into the program by the linker:

void
_init(void)
{
	extern u32int bdata[], edata[], end[], etext[];
	u32int *src, *dst;
	src = etext;
	dst = bdata;

	while(dst < (u32int *)edata)
		*dst++ = *src++;
	while(dst < (u32int *)end)
		*dst++ = 0;
}

With global variables in place, I can redefine my registers as values instead of preprocessor macros:

RCCreg	const *RCC	= (void*)0x40021000;
GPIOreg	const *GPIOA	= (void*)0x48000000;
GPIOreg	const *GPIOB	= (void*)0x48000400;
GPIOreg	const *GPIOC	= (void*)0x48000800;
SCBreg	const *SCB	= (void*)0xE000ED00;
STKreg	const *STK	= (void*)0xE000E010;

While this arguably wastes memory and might even generate worse code, I do it anyway, because it makes those symbols available to me in the debugger, and substitutes the address they point to with their name in assembly output, making it easier to read. With the appropriate types defined in an include file for acid, I can do things like this from the debugger:

acid: RCC->ahb2enr\B
0b00000000000000000000000000000011

The waiting game

Factotum is going to spend 99.999% of its time doing nothing; waiting for input, waiting for a button press, waiting for a timeout, waiting for data to be comitted to flash, waiting for its turn to transmit data. It will even wait for multiple things at once. So I'd better find a good way to wait.

In my blinking light demo, I used a busy loop to create time between blinks:

m	= 0x00080008;
r	= 0x00000008;
while(1){
	GPIOB->bsr = r;
	r  ^= m;
	for(i = 0; i < 800000; i++);
}

If the CPU has nothing useful to do while it's waiting, I would like it to go to sleep and use less power.

There are quite a few low-power states available for this chip, with various implications about what components remain operational. For now, I am going to choose the wait for interrupt instruction, as it should be available on most ARM processors, and does exactly what I need; go to sleep until an interrupt fires.

This instruction is not implemented in the existing toolchain. But simple instructions like this can be implemented as a macro:

#define WFI WORD $0xbf00bf30
TEXT	wfi(SB), 4, $-4
	WFI
	RET

As an aside, I discussed adding this instruction and the other 3 "hint" instructions (YIELD, WFE, SEV) to the assembler and linker with the core 9front developers. The consensus was that it is not worth it to add an instruction of no arguments, which is really only used for kernels, to the object files and libraries, which may be parsed by older linkers. If the instruction were used more frequently or had elaborate syntax or arguments, it may be worth implementing in the assembler, and have it emit opaque bytes that the linker doesn't have to understand. But for these, the tradeoff isn't worth it.

I can utilize the SysTick timer to increment a global variable once per millisecond. It should also have the side effect of waking up the CPU. Here is my first attempt at a sleep function:

extern void wfi(void);
static int msticks;

/* this goes into vtable slot 15 for SysTick interrupt */
void systick(void) { msticks++; }

void
sleep(int ms)
{
	long start = msticks;
	do {
		wfi();
	} while(msticks - start < ms);
}

But this won't do anything until I configure the SysTick timer and give it a clock source.

Clocks. Clocks everywhere

There are so many clocks and timers on this device. Some clock sources can even be calibrated by other clock sources, or feed into a frequency multiplier. There can be different clocks used for the CPU, and each of the peripheral buses, and in some cases, even the peripherals themselves. It's all pretty overwhelming. Luckily there is a pretty gentle introduction to the clocks on ST's site[1].

For the time being, I am still at the "blink a LED" stage. The manual for this device states the following:

The MSI is used as system clock source after startup from Reset, configured at 4 MHz.

The Multi-speed internal (MSI) RC oscillator, as the name implies, is special because it is adjustable; it can be run from 100Khz all the way up to 48 MHz, and can be scaled up and down on the fly to save power or improve performance.

I am just going to try using the default clock at 4 MHz. I will go through the motions of configuring it as an exercise, but the settings I configure should reflect the default settings after a reset:

enum {
	SW	= 0b00000011,
	SWS	= 0b00001100,
	MSIRNG	= 0b11110000,

	MSISW	= 0b00	<< 0,
	MSISWS	= 0b00	<< 2,
	MSI4Mhz	= 0b0110<< 4,
};
while((RCC->cr & MSIRDY) == 0)
	RCC->cr |= MSION;
while((RCC->cr & MSIRNG) != MSI4Mhz)
	RCC->cr |= MSI4Mhz;
while((RCC->cfgr & SWS) != MSISWS)
	RCC->cfgr = (RCC->cfgr & ~SW) | MSISW;

This may look a bit silly because the MSI bit patterns are just 0, so the code to enable them could be simpler. But I'm writing it with the expectation that I will tweak it in the future. With the system clock configured to run at 4Mhz, I need to load the SysTick timer appropriately:

enum {
	Hz = 4000000,
};
STK->load = Hz / 1000;
while((STK->ctrl & 0b111) != 0b111)
	STK->ctrl = 0b111;

Per the manual[2], setting bit 0 in the STK_CTRL register enables the timer, setting bit 1 tells it to generate an interrupt when it reaches 0, and bit 2 chooses between using the system clock (1) or the system clock divided by 8.

The main loop now looks like this:

while(1){
	r ^= m;
	GPIOB->bsr = sr;
	sleep(1000);
}

And, after some debugging, this works! The documentation states that the MSI has more drift than the other clocks, especially at extreme temperatures, but for this demo it suffices.

Here is the state of the repository at this point. Clock configuration will become more important when I start setting up USB, which is sensitive to clock inaccuracies.

UART console

Although I have a debugger, and it has helped me get this far, the device runs differently under a debugger, and may not match "real" usage. USB transactions have strict time limits, counted in milliseconds. I will not be able to stop execution and step through the system's state without those transactions failing. What I really want is a way to emit debug logs that I can read from my workstation. To do that, I got a USB serial cable adapter and connected it to pins D1 (PA9), D2 (PA10), and GND of my development board.

I left power disconnected, as I am still powering the device through the onboard ST-link connected to the micro USB port. The adapter forwards the USB 5V power, which I should be able to use to power the device through the 5V pin, but I have read that some versions of this development board will prevent the device from starting up in this configuration without extra steps, which I'm not interested in taking right now.

The USART peripheral has a ton of features; section 39 of the reference manual[3] alone is 68 pages. For my purposes, almost none of these features are needed, and we can use the peripheral with its default settings. It was helpful to get a really dumb solution working first, to figure out what needed to be initialized and what didn't:

int i;
char msg[] = "Hello, world!\r\n";
u32int blink = 0x00080000;
for(;;){
	GPIOB->bsr = (blink ^= 0x00080008);
	for(i = 0; i < sizeof(msg)-1; i++){
		while((USART1->isr & TXE) == 0)
			;
		USART1->tdr = msg[i];
	}
	sleep(1000);
}

After a lot of trial and error, mostly around configuring the GPIO pins, I hit my first milestone:

$ tio /dev/ttyUSB0
[12:39:56.743] tio 3.9
[12:39:56.743] Press ctrl-t q to quit
[12:39:56.756] Connected to /dev/ttyUSB0
Hello, world!
Hello, world!
Hello, world!

Once more with feeling DMA

For my purposes, I want similar semantics to a pipe on a Unix or Plan 9 system, with a focus on higher throughput. I intend to use this for printf-style debugging, so I want the USART to remain responsive even when the system is busy. That means configuring the Direct Memory Access controller (DMA) so I can transmit and receive data while my program is still running.

My next goal is to implement an "echo" server; read data from RX line and send it over the TX line. There will be no message delimiter; whatever is read will be sent out ~immediately.

Ring buffers

Everyone should be familiar with ring buffers, and they are a crucial tool for streaming an unbounded flow of data with a fixed amount of memory. I will use them here, and it's useful to introduce them separately:

typedef struct Ring Ring;
struct Ring
{
	int head;
	int tail;
	int mask;
};

static char* ptr(Ring *v, int x) {
	return (char*)v + sizeof(*v) + (x & v->mask);
}

int length(Ring *v) { return v->tail - v->head; }
int full(Ring *v) { return length(v) == v->mask + 1; }
int empty(Ring *v) { return v->head == v->tail; }

The Ring struct is meant to be embedded in another struct and immediately followed by the byte buffer where data is stored, like so:

struct {
	Ring;
	uchar buf[64];
} tx, rx;
tx.mask = sizeof(tx.buf)-1;
rx.mask = sizeof(rx.buf)-1;

It is probably more common to see ring buffers where each slot is a pointer to a location in memory. That is generally easier to work with, as you can update each slot atomically, you can preserve the size of each message, and you don't have to deal with a message being split across the edge of the buffer. However, the format above is easier to allocate statically, and I am trying to avoid dynamic memory management as long as I can.

I am making use of a very common trick; by making the buffer size a power of 2, I can allow the head and tail counters to overflow freely, without needing to reset them when they go beyond the buffer area. Copying data from one ring to another then looks like this:

int
copy(Ring *dst, Ring *src)
{
	int n;
	
	for(n = 0; !full(dst) && !empty(src); n++)
		*ptr(dst, dst->tail++) = *ptr(src, src->head++);
	return n;
}

As we will see later, it is also useful to have a function that outlines the largest contiguous chunk of memory at the front of the buffer:

typedef struct Seg Seg;
struct Seg {
	void *addr;
	int len;
};

Seg
next(Ring *v)
{
	int len, edge;

	edge = (v->head | v->mask) + 1;
	len = edge - v->head;

	if(len > length(v))
		len = length(v);

	return (Seg){ptr(v, v->head), len};
}

with the ring buffer in place, we can try configuring the DMA controller to read from it and write into it.

Transmit

Transmission is relatively easy; because we know how much data we have to send, we can just fill the transmit buffer and trigger a DMA.

inflight.len = 0;
for(;;){
	if(txch->cndtr == 0 && length(&tx) - inflight.len > 0){
		tx.head += inflight.len;
		inflight = next(&tx);
		txch->ccr &= ~EN;
		txch->cmar = (u32int)inflight.addr;
		txch->cndtr = inflight.len;
		txch->ccr |= EN;
	}
}

Receive

To receive data, because we don't know how much data is coming, I will put the DMA into "circular buffer" mode, where it will wrap around and write into the beginning of the buffer once it reaches the end. Then I can chase the DMA as it fills the receive buffer:


for(;;){
	rxpos = sizeof(rx.buf) - rxch->cndtr;
	rx.tail += (rxpos - rx.tail) & rx.mask;
	copy(&tx, &rx);
	if(txch->cndtr == 0 && length(&tx) > 0){
		txch->ccr &= ~EN;
		txch->cndtr = next(&tx, (uintptr*)&txch->cmar, sizeof tx.buf);
		txch->ccr |= EN;
	}
}

receive buffer overrun and flow control

The DMA's circular mode is a little bit limited in that the count of pending transfers is reset whenever a wrap-around occurs. It would have been better if this register just decremented forever, and required the software to mask or modulo it. As it stands, there is no way to reliably detect how much the buffer has been overrun. Consider the scenario where we have a buffer of 16 bytes. H is the position of the head, T the tail, and D is the next position the DMA will read into. For whatever reason, the CPU cannot keep up with the DMA:

T0 W H A T . I S . S I H T D T1 T . I S . S I X . T I M E H T D T2 S . S E V E S . S I X . T I M E H T D T3 S . S E V E N ? S I X . T I M E H T D

You can see at time T2, the DMA has overwritten the head of the ring buffer. The only thing we can do at this point is to advance the tail to the last known location the DMA wrote into, and lose all of the data between the old tail and head. The overflow condition itself can be detected whenever the tail crosses over the head. This information is actually implicitly stored in the high bits of the head and tail positions which are normally masked off. The number of overruns can be recovered like so:

while(length(&rx) > sizeof rx.buf){
	rx.head += sizeof rx.buf;
	overruns++;
}

For this use case, I am okay with losing data on overruns. My thinking is, for a baud rate of 115200 and a clock speed of 32Mhz, I have well over 200 clock cycles to stay ahead of the DMA. I could send software flow control signals (XON/XOFF) before the receive buffer gets full, but I would still have to consider what would happen if the flow control signal doesn't reach the peer in time, or if the peer doesn't respect it. In the future, there will be cases where I want to prioritize new data by overwriting old, and cases where I want to preserve old data by applying back pressure.

How to wait

I currently have a busy loop:

for(;;){
	rxpos = sizeof(rx.buf) - rxch->cndtr;
	rx.tail += (rxpos - rx.tail) & rx.mask;
	copy(&tx, &rx);
	if(txch->cndtr == 0 && length(&tx) - inflight.len > 0){
		tx.head += inflight.len;
		inflight = next(&tx, sizeof tx.buf);
		txch->ccr &= ~EN;
		txch->cmar = (u32int)inflight.addr;
		txch->cndtr = inflight.len;
		txch->ccr |= EN;
	}
}

If there is no data to send or receive, the loop will spin forever, polling DMA registers. I have gone through all this trouble to offload the movement of data, only to waste the CPU's time polling for completion. It's fine for a demo, but in real usage, when there is no data to read or write, I want my CPU to do other work, or go to sleep to save power. I can put the CPU to sleep when there is no data to send or receive by putting this in my loop.

	do {
		rxpos = sizeof(rx.buf) - rxch->cndtr;
		if(((rxpos - rx.tail) & rx.mask) > 0)
			break;
		if(txch->cndtr == 0 && inflight.len > 0)
			break;
		sleep(0);
	}while(1);

The sleep function, even if given a zero duration, has the side effect of executing the WFI instruction, which will put the CPU to sleep until an interrupt fires. I have to enable these interrupts.

The interrupt handlers themselves don't have to do anything other than clear the interrupt status registers. For a demo, they can be as simple as:

void dma1irq(void)	{ DMA1->ifcr = ~0; }
void usart1irq(void)	{ USART1->icr = ~0; }

Although in a "real" program, they should check for error interrupts and do something about them, like reset the peripheral or something.

I went a little overboard in implementing this because I want to use this as preparation for building a more general-purpose pipe abstraction. They will be used not just for printing debug logs, but as the foundation for implementing USB pipes. It may also come in handy when reading data from flash or other peripherals. If I can trip over all the mines at this early stage, it should be easier much later.

The code at this point is here.

  1. Introduction to the STM32 microcontroller clock system

    ↩︎︎
  2. STM32 Cortex®-M4 MCUs and MPUs programming manual

    ↩︎︎
  3. STM32L41xxx/42xxx/43xxx/44xxx/45xxx/46xxx advanced Arm® -based 32-bit MCUs

    ↩︎︎