Learning C
C is a general purpose programming language which was created in 1972 by Dennis Ritchie.
While it was originally designed to implement operating systems, the features which made the language ideal for that purpose, also make it ideal for developing software for small resource constrained MCUs.
On this page we will go through how to use C for embedded development.
Hello World - Desktop vs. MCU
When learning any programming language, the first "project" will typically be a small program which prints out "Hello World!". C was designed to give programmers relatively direct access to the target hardware. This was ideal for creating operating systems and indeed most, if not all, serious operating systems are largely implemented in C. This design also make C an ideal language for developing embedded software targeting small MCUs with very limited resources.
Standard UNIX
On a standard UNIX system, that program, implemented in standard C, would typically look something like this:
1 #include <stdio.h>
2
3 int main(int argc, char** argv) {
4 printf("Hello World!\n");
5 return 0;
6 }
Let us break down the classic C program line by line.
Line 1 contains a Preprocessor Directive. This essentially instruct the compiler to include a header file called stdio.h (Standard Input/Output). This file in turn contains the blueprint (called "function prototype") for functions interacting with the I/O system in the operating system, such as printf.
In line 3 follows a function declaration, creating a function named main. Every C program includes a function called main and this function serves as the entry point. The int tell the compiler that this function will return an integer value and the int argc is a variable containing how many command-line arguments was received, and char** argv is an array (think of a list) with those arguments). The actual function code is enclosed in a code block using curly brackets.
In line 4 we call the function printf which will print a string on the terminal executing the program. The printf is provided with an argument, which is a string enclosed in double-quotes "Hello World!". The compiler will include this string as a constant.
Finally in line 5, we return the value 0. The convention is that returning a non-zero value indicates an error.
Most UNIX like systems have (or bloody should have) a C compiler, typically named cc, so above hello.c can be built into a binary executable thus:
$ cc hello.c -o hello
This will produce a binary hello, which can now be executed:
$ ./hello
Hello World!
Embedded
On an embedded system using any MCU whether it be STM32, RP2350 or any other, things get a little bit more complicated. On a UNIX system, the kernel knows how to execute a binary, and the shell (for example Bash) know how to load it, pass parameters and interpret the return value. Typically, a MCU core will simply jump to a specific address - often 0x00000000. The actual format of the binary depends on the MCU core in use and the architecture. There are seldom a general way to print information, so typically the developer will need to specify how to print. Finally, an embedded application will never actually exit, so the main() will typically run an infinite loop (called the "super loop" or "main loop"), or hand over control to an operating system that doesn't return.
During boot, typically, part of the MCU will need to be initialized. It is possible to do this manually, but it will often be handled by some sdk provided by the vendor of the MCU. This startup code will typically lead up to calling main().
Skipping the actual MCU initialization and only blinking a LED rather than printing a statement, a very minimal main() could look something like this:
1 int main(void) {
2 uint16_t led = PIN('C', 13); // Blue LED
3 RCC->AHB1ENR |= BIT(PINBANK(led)); // Enable GPIO clock for LED
4 gpio_set_mode(led, GPIO_MODE_OUTPUT); // Set blue LED to output mode
5 for (;;) {
6 gpio_write(led, true);
7 spin(999999);
8 gpio_write(led, false);
9 spin(999999);
10 }
11 return 0;
12 }
This will all appear quite incomprehensible for a C beginner, but for now let's just realize that a C program always have a main() function and on embedded systems this function will normally never exit since there is nowhere else to go.
Variables & Fixed-Width Types
Variables is one of the key concepts in (almost) any programming language. A variable is a named container which can contain any value of the type of the variable.
Variable declaration
The most basic variable declaration in C looks like this:
int i = 10;
This code snippet will declare (create) a variable by the name i with a type of int and assigning an initial value of 10.
The type int is an integer. The actual number of bits in an int depends on the platform but on a 32 bit MCU it is most often 32 bits. The value is signed, so it can hold values in the following range:
-2,147,483,648 to 2,147,483,647 (2^31 − 1)
Variable types
C only defines a few basic types:
| Type | Range | Description |
|---|---|---|
| char | 0 - 255 | A single byte capable of holding one character |
| int | -2,147,483,648 - 2,147,483,647 | An integer - on modern 32-bit MCUs typically a 32-bit value |
| float | ±1.175 × 10⁻³⁸ to ±3.403 × 10³⁸ | 32-bit floating point value |
| double | ±1.7 × 10^308 | 64-bit (8 byte) floating point |
The types can be modified by applying a so-called modifier, which can be either long or short, and for int types it can also be unsigned.
long int i = 10;
The compiler will also allow various emissions and stacking.
short i = 10;
long i = 10;
unsigned short i = 10;
unsigned long long int i = 10;
Fixed width types
In K&R, the following is mentioned:
Each compiler is free to choose appropriate sizes for its own hardware, subject only to the the restriction that shorts and ints are at least 16 bits, longs are at least 32 bits, and short is no longer than int, which is no longer than long
This unpredictable nature can be troublesome when creating libraries that need to work on multiple platforms, or used for protocol and data format implementations. For that reason, a library called stdint.h was added in C99 containing: uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t for unsigned and signed integer types of specific number of bits.
Example
Consider the following code running on a STM32 MCU:
1 printf("\n\n\n--------\nStarting\n");
2
3 printf("sizeof(char) = %d\n", sizeof(char));
4 printf("sizeof(short int) = %d\n", sizeof(short int));
5 printf("sizeof(int) = %d\n", sizeof(int));
6 printf("sizeof(long int) = %d\n", sizeof(long int));
7 printf("sizeof(long long int) = %d\n", sizeof(long long int));
8 printf("sizeof(float) = %d\n", sizeof(float));
9 printf("sizeof(double) = %d\n", sizeof(double));
10 printf("sizeof(long double) = %d\n", sizeof(long double));
11 printf("sizeof(uint8_t) = %d\n", sizeof(uint8_t));
12 printf("sizeof(uint16_t) = %d\n", sizeof(uint16_t));
13 printf("sizeof(uint32_t) = %d\n", sizeof(uint32_t));
14 printf("sizeof(uint64_t) = %d\n", sizeof(uint64_t));
15 printf("sizeof(void *) = %d\n", sizeof(void*));
This will produce the following output:
-------- Starting sizeof(char) = 1 sizeof(short int) = 2 sizeof(int) = 4 sizeof(long int) = 4 sizeof(long long int) = 8 sizeof(float) = 4 sizeof(double) = 8 sizeof(long double) = 8 sizeof(uint8_t) = 1 sizeof(uint16_t) = 2 sizeof(uint32_t) = 4 sizeof(uint64_t) = 8 sizeof(void *) = 4
Running the same on a Debian 64-bit OS,
-------- Starting sizeof(char) = 1 sizeof(short int) = 2 sizeof(int) = 4 sizeof(long int) = 8 sizeof(long long int) = 8 sizeof(float) = 4 sizeof(double) = 8 sizeof(long double) = 16 sizeof(uint8_t) = 1 sizeof(uint16_t) = 2 sizeof(uint32_t) = 4 sizeof(uint64_t) = 8 sizeof(void *) = 8
Notice the difference of sizes of int, long int and others, but the same sizes of uint8_t, uint16_t and so forth.
Type Casting
Scope
Variables in C can be declared in different locations (scopes). If a variable is declared before any function declarations they are said to be global, but a variable can also be declared inside a function which will limit the scope to that particular functions. For this reason, local variables can have the same name and each will refer to the local value.
Expressions
Operators
Precedence
Control Flow & Decision Making
In C there are multiple ways to control the flow through the program.
The if-else
1 if (x == 10) {
2 // Do something
3 } else {
4 // Do something else
5 }
Loops & Iteration
The Super Loop
As mentioned earlier, embedded programs are never supposed to exit, since they got nowhere to go really. They often achieve this by running an endless loop, known as the super loop:
1 while (1) { // Loop forever
2 // Do something here
3 }
The endless loop can be written in multiple ways. Another often seen approach is:
1 for (;;) { // Loop forever
2 // Do something here
3 }
Which one is the most efficient is a topic of an almost religious discussion. There might have been a time where one was better than the other (better as in wasting less cpu cycles) but with modern optimizing compilers I doubt very much that there will be any difference.
Fans of "The Hitchhikers Guide to the Galaxy" will appreciate:
1 while (42) { // Loop forever
2 // Do something here
3 }
This will execute exactly like the first one, taking advantage of the fact that anything non-zero is "true" per definition.
For the final approach we'll look at a feature seldom used in C programs. K&R describes:
C provides the infinitely-abusable goto statement, and labels to branch to. Formally, the goto statement is never necessary, and in practice it is almost always easy to write code without it.
Apart from being hilarious it does open the possibility of doing:
1 loop_start:
2 //Do something
3 goto loop_start;
While I have never seen this in the wild, it is technically valid C syntax. One reason it is quite easy to end up with the worst spaghetti code possible, as the compiler will not guard against structural mismatching.
Functions & Modular Code
To be added
Bitwise Operations Masterclass
To be added
Arrays
To be added
Pointers Demystified
To be added
Pointer Arithmetic
To be added
Structs & Memory Alignment
To be added
Unions
To be added
Function Pointers & Callbacks
To be added
The volatile Keyword
To be added
Enums & State Machines
To be added
Dynamic Memory vs. Static Allocation
To be added
Macros, Pragmas & Preprocessor
To be added
Common Compile Errors
To be added