文章

[C++并发编程] 线程生命周期与RAII

[C++并发编程] 线程生命周期与RAII

std::thread基础

std::thread构造

std::thread的构造函数接受一个可调用对象以及可选的参数列表,可以通过三种方式构造:

  • 函数指针
  • Lambda表达式
  • 函数对象

最常用的就是通过函数指针和Lambda表达式进行构造。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
#include <thread>
#include <iostream>
#include <vector>

void print_hello(int id)
{
    std::cout << "Hello from thread " << id << "\n";
}

int main()
{
    std::vector<int> data = {1, 2, 3, 4, 5};
    int sum = 0;
    std::thread t(print_hello, 42);	//函数指针
    std::thread t1([&data, &sum]() {	//Lambda表达式
        for (int v : data) {
            sum += v;
        }
    });
    
    t.join();
    t1.join();
    
    return 0;
}

join()和detach()

join()为阻塞调用,当前线程会被阻塞,直到目标线程执行完毕才继续往下走:

detach()会把线程从std::thread对象的管理中剥离出来,剥离后的线程会在后台独立运行,std::thread对象不再持有任何对它的引用,我们无法再join它了。

如果在一个joinable()的std::thread对象上既不调用join()也不调用detach(),线程对象在析构时会直接崩溃。所以我们必须在每一条代码路径上处理线程的join/detach,包括异常路径。一个常见的模式是使用RAII包装器,构造时保存线程,析构时自动join。

最基本的线程使用模式就是为每个子任务派生一个线程,在当前作用域退出时join所有线程:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
#include <thread>
#include <iostream>
#include <vector>
#include <chrono>
#include <algorithm>

void process_range(const std::vector<int>& input,
                   std::vector<int>& output,
                   std::size_t start,
                   std::size_t end)
{
    for (std::size_t i = start; i < end; ++i) {
        // 模拟一个计算密集型操作
        output[i] = input[i] * input[i];
    }
}

int main()
{
    constexpr std::size_t kDataSize = 10'000'000;
    constexpr unsigned int kNumThreads = 4;

    std::vector<int> input(kDataSize);
    std::vector<int> output(kDataSize);

    // 初始化输入数据
    for (std::size_t i = 0; i < kDataSize; ++i) {
        input[i] = static_cast<int>(i);
    }

    auto start_time = std::chrono::high_resolution_clock::now();

    std::vector<std::thread> threads;
    threads.reserve(kNumThreads);

    std::size_t chunk_size = kDataSize / kNumThreads;

    // 派生线程
    for (unsigned int i = 0; i < kNumThreads; ++i) {
        std::size_t start = i * chunk_size;
        std::size_t end = (i == kNumThreads - 1)
                              ? kDataSize
                              : start + chunk_size;
        threads.emplace_back(process_range,
                             std::cref(input),
                             std::ref(output),
                             start,
                             end);
    }

    // 在作用域退出前 join 所有线程
    for (auto& t : threads) {
        t.join();
    }

    auto end_time = std::chrono::high_resolution_clock::now();
    auto ms = std::chrono::duration_cast<std::chrono::milliseconds>(
                  end_time - start_time);

    std::cout << "Processed " << kDataSize << " elements in "
              << ms.count() << " ms using "
              << kNumThreads << " threads\n";
    return 0;
}

线程的标识与查询

get_id()

每个线程都有一个唯一的标识符,类型是 std::thread::id。你可以通过 std::thread::get_id() 获取某个线程对象的 ID,也可以通过 std::this_thread::get_id() 获取当前线程的 ID。

hardware_concurrency()

std::thread::hardware_concurrency() 是一个静态成员函数,返回一个提示值,表示当前系统上真正可以并发执行的线程数量。如果信息不可用,函数返回0。

线程函数的异常

异常永远不应该逃逸线程函数。如果一个异常从线程函数中逃逸出去(即线程函数抛出了异常但没有在线程内部被 catch),std::terminate() 会被调用,程序直接崩溃。因为每个线程都有自己独立的调用栈,异常处理机制只能在当前线程的栈上工作,如果异常穿透了线程函数,那意味着没有一个catch块能接住这个异常。

decay-copy(退化拷贝)

std::thread的构造函数会按值拷贝或移动所有传入的参数:引用被剥离、const/volatile被丢弃、数组退化为指针、函数退化为函数指针。

decay-copy的设计动机就是为了让每个线程都默认拥有自己的一份参数副本,避免隐式的共享状态。如果我们要共享,必须使用std::ref/std::cref引用包装器,线程会拷贝引用包装器,引用包装器内部的指针会指向原变量,从而实现共享。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
#include <thread>
#include <iostream>
#include <functional>
#include <string>

void append_suffix(std::string& str, const std::string& suffix)
{
    str += suffix;
}

int main()
{
    std::string message = "Hello";
    std::string suffix = " World";

    std::thread t(append_suffix, std::ref(message), std::cref(suffix));
    t.join();

    std::cout << message << "\n";  // 输出 "Hello World"
    return 0;
}

线程的生命周期与引用对象的生命周期

有一大类并发bug的来源都是因为:线程的生命周期超过了它所引用对象的生命周期。解决方案就是复制到线程,或用shared_ptr延长生命周期。更好的方案就是在绝大多数场景下不要使用detach,使用join和RAII来配合。

线程所有权与RAII

std::thread 是不可复制的。你不能把一个线程对象赋值给另一个,也不能通过值传递来转移它。所以 std::thread 只支持 move 语义。

joining_thread:接管所有权的 RAII wrapper

wrapper 拥有线程,wrapper 析构时自动 join。这个思路的实现就是 joining_thread,它本质上就是 C++20 std::jthread

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
#include <thread>
#include <utility>

class JoiningThread {
public:
    JoiningThread() noexcept = default;

    // 接受任意可调用对象和参数,直接构造线程
    template <typename Callable, typename... Args>
    explicit JoiningThread(Callable&& func, Args&&... args)
        : thread_(std::forward<Callable>(func), std::forward<Args>(args)...)
    {}

    // 从 std::thread move 构造——接管所有权
    explicit JoiningThread(std::thread t) noexcept
        : thread_(std::move(t))
    {}

    // 支持从另一个 JoiningThread move
    JoiningThread(JoiningThread&& other) noexcept
        : thread_(std::move(other.thread_))
    {}

    JoiningThread& operator=(JoiningThread&& other) noexcept
    {
        if (this != &other) {
            // 先处理当前持有的线程
            if (joinable()) {
                join();
            }
            thread_ = std::move(other.thread_);
        }
        return *this;
    }

    // 也可以从一个新的 std::thread 赋值
    JoiningThread& operator=(std::thread other) noexcept
    {
        if (joinable()) {
            join();
        }
        thread_ = std::move(other);
        return *this;
    }

    ~JoiningThread()
    {
        if (joinable()) {
            join();
        }
    }

    void join()
    {
        thread_.join();
    }

    void detach()
    {
        thread_.detach();
    }

    bool joinable() const noexcept
    {
        return thread_.joinable();
    }

    // 获取底层 std::thread(用于 native_handle 等)
    std::thread& get() noexcept { return thread_; }
    const std::thread& get() const noexcept { return thread_; }

    // 禁止复制
    JoiningThread(const JoiningThread&) = delete;
    JoiningThread& operator=(const JoiningThread&) = delete;

private:
    std::thread thread_;
};

虽然我们实现了一个自动join的RAII 包装器,但是join()本身也是会抛出异常的,所以需要在析构函数中用try-catch把join()包起来。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
~JoiningThread()
{
    if (joinable()) {
        try {
            join();
        }
        catch (const std::system_error& e) {
            // join 失败了,记录日志但不抛出
            // 在实际项目中应该用正式的日志系统
            std::fprintf(stderr,
                         "JoiningThread: join() failed: %s\n", e.what());
        }
    }
}

C++20 的 std::jthread

C++20 标准终于引入了 std::jthread,它的行为跟我们的 JoiningThread 非常相似——析构时自动 join。但 std::jthread 还多了一个重要的功能:协作式取消(cooperative cancellation),它内部持有一个 std::stop_source,可以通过 request_stop() 请求线程停止执行。

thread_local

thread_local 就是线程存储持续时间的说明符——被它修饰的变量在每个线程中都有自己独立的实例,线程从创建到退出期间一直存在。

thread_local 变量的初始化发生在每个线程首次使用(ODR-use)时,而不是程序启动时。这个”首次使用时初始化”的行为非常重要——它保证了以下几点:第一,如果一个 thread_local 变量从来没有被某个线程访问过,那个线程就不会为它分配内存或执行初始化,所以不会有浪费。第二,初始化是线程安全的——标准保证即使多个线程同时首次访问同一个 thread_local 变量,每个线程的初始化也只会执行一次,而且不会互相干扰。

这个”延迟初始化”的特性让 thread_local 非常适合实现一些”按需分配”的资源——比如每个线程的随机数生成器、内存池、日志缓冲区等。这些资源如果全局共享就需要加锁,而用 thread_local 之后就完全无锁了,下面是一个thread_local 在性能优化中的典型用法:给每个线程一个小的内存池,小对象的分配直接从线程本地池中取,不用跟其他线程竞争

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
#include <vector>
#include <cstddef>

class ThreadLocalPool {
public:
    static ThreadLocalPool& instance()
    {
        thread_local ThreadLocalPool pool;
        return pool;
    }

    void* allocate(std::size_t size)
    {
        if (size <= kBlockSize) {
            if (!free_list_.empty()) {
                void* ptr = free_list_.back();
                free_list_.pop_back();
                return ptr;
            }
            // 从大块中切出一块
            if (current_offset_ + size > kChunkSize) {
                chunks_.emplace_back(new char[kChunkSize]);
                current_offset_ = 0;
            }
            void* ptr = chunks_.back().get() + current_offset_;
            current_offset_ += size;
            return ptr;
        }
        // 超过块大小的分配,回退到全局分配器
        return ::operator new(size);
    }

    void deallocate(void* ptr, std::size_t size)
    {
        if (size <= kBlockSize) {
            free_list_.push_back(ptr);
        }
        else {
            ::operator delete(ptr);
        }
    }

private:
    ThreadLocalPool() = default;

    static constexpr std::size_t kBlockSize = 256;
    static constexpr std::size_t kChunkSize = 4096;

    std::vector<std::unique_ptr<char[]>> chunks_;
    std::vector<void*> free_list_;
    std::size_t current_offset_{kChunkSize};  // 初始值触发首次分配
};

std::call_once与std::once_flag

std::call_once 是 C++11 提供的一次性初始化机制。你给它一个 std::once_flag 和一个可调用对象,它保证无论有多少个线程同时调用 call_once,可调用对象只被执行一次——第一个到达的线程执行初始化,其余线程等待它完成。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
#include <mutex>
#include <iostream>
#include <thread>

std::once_flag init_flag;
int* shared_resource = nullptr;

void ensure_initialized()
{
    std::call_once(init_flag, []() {
        std::cout << "Initializing shared resource...\n";
        shared_resource = new int(42);
    });
}

void use_resource(const char* thread_name)
{
    ensure_initialized();
    std::cout << thread_name << ": resource = " << *shared_resource << "\n";
}

int main()
{
    std::thread t1(use_resource, "Thread-A");
    std::thread t2(use_resource, "Thread-B");
    std::thread t3(use_resource, "Thread-C");

    t1.join();
    t2.join();
    t3.join();

    delete shared_resource;
    return 0;
}

输出中会发现 “Initializing shared resource…” 只出现了一次——无论三个线程的调度顺序如何,初始化代码只执行了一次。std::once_flag 记录了初始化是否已经完成,call_once 在每次调用时检查这个标志位。如果初始化还没开始,第一个线程执行初始化;如果正在进行中,其他线程阻塞等待;如果已经完成,所有线程直接跳过。

std::call_once 有一个很关键的行为:如果初始化函数(可调用对象)抛出了异常,call_once 不会把 once_flag 标记为”已完成”。这意味着下一次有线程调用 call_once 时,初始化会再次尝试。

从 C++11 开始,函数作用域中的 static 局部变量有一个非常重要的保证:它的初始化是线程安全的。如果有多个线程同时首次执行到 static 变量的声明处,只有一个线程会执行初始化,其他线程会等待。

本文由作者按照 CC BY 4.0 进行授权