[C++并发编程] 线程生命周期与RAII
std::thread基础
std::thread构造
std::thread的构造函数接受一个可调用对象以及可选的参数列表,可以通过三种方式构造:
- 函数指针
- Lambda表达式
- 函数对象
最常用的就是通过函数指针和Lambda表达式进行构造。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
#include <thread>
#include <iostream>
#include <vector>
void print_hello(int id)
{
std::cout << "Hello from thread " << id << "\n";
}
int main()
{
std::vector<int> data = {1, 2, 3, 4, 5};
int sum = 0;
std::thread t(print_hello, 42); //函数指针
std::thread t1([&data, &sum]() { //Lambda表达式
for (int v : data) {
sum += v;
}
});
t.join();
t1.join();
return 0;
}
join()和detach()
join()为阻塞调用,当前线程会被阻塞,直到目标线程执行完毕才继续往下走:
detach()会把线程从std::thread对象的管理中剥离出来,剥离后的线程会在后台独立运行,std::thread对象不再持有任何对它的引用,我们无法再join它了。
如果在一个joinable()的std::thread对象上既不调用join()也不调用detach(),线程对象在析构时会直接崩溃。所以我们必须在每一条代码路径上处理线程的join/detach,包括异常路径。一个常见的模式是使用RAII包装器,构造时保存线程,析构时自动join。
最基本的线程使用模式就是为每个子任务派生一个线程,在当前作用域退出时join所有线程:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
#include <thread>
#include <iostream>
#include <vector>
#include <chrono>
#include <algorithm>
void process_range(const std::vector<int>& input,
std::vector<int>& output,
std::size_t start,
std::size_t end)
{
for (std::size_t i = start; i < end; ++i) {
// 模拟一个计算密集型操作
output[i] = input[i] * input[i];
}
}
int main()
{
constexpr std::size_t kDataSize = 10'000'000;
constexpr unsigned int kNumThreads = 4;
std::vector<int> input(kDataSize);
std::vector<int> output(kDataSize);
// 初始化输入数据
for (std::size_t i = 0; i < kDataSize; ++i) {
input[i] = static_cast<int>(i);
}
auto start_time = std::chrono::high_resolution_clock::now();
std::vector<std::thread> threads;
threads.reserve(kNumThreads);
std::size_t chunk_size = kDataSize / kNumThreads;
// 派生线程
for (unsigned int i = 0; i < kNumThreads; ++i) {
std::size_t start = i * chunk_size;
std::size_t end = (i == kNumThreads - 1)
? kDataSize
: start + chunk_size;
threads.emplace_back(process_range,
std::cref(input),
std::ref(output),
start,
end);
}
// 在作用域退出前 join 所有线程
for (auto& t : threads) {
t.join();
}
auto end_time = std::chrono::high_resolution_clock::now();
auto ms = std::chrono::duration_cast<std::chrono::milliseconds>(
end_time - start_time);
std::cout << "Processed " << kDataSize << " elements in "
<< ms.count() << " ms using "
<< kNumThreads << " threads\n";
return 0;
}
线程的标识与查询
get_id()
每个线程都有一个唯一的标识符,类型是 std::thread::id。你可以通过 std::thread::get_id() 获取某个线程对象的 ID,也可以通过 std::this_thread::get_id() 获取当前线程的 ID。
hardware_concurrency()
std::thread::hardware_concurrency() 是一个静态成员函数,返回一个提示值,表示当前系统上真正可以并发执行的线程数量。如果信息不可用,函数返回0。
线程函数的异常
异常永远不应该逃逸线程函数。如果一个异常从线程函数中逃逸出去(即线程函数抛出了异常但没有在线程内部被 catch),std::terminate() 会被调用,程序直接崩溃。因为每个线程都有自己独立的调用栈,异常处理机制只能在当前线程的栈上工作,如果异常穿透了线程函数,那意味着没有一个catch块能接住这个异常。
decay-copy(退化拷贝)
std::thread的构造函数会按值拷贝或移动所有传入的参数:引用被剥离、const/volatile被丢弃、数组退化为指针、函数退化为函数指针。
decay-copy的设计动机就是为了让每个线程都默认拥有自己的一份参数副本,避免隐式的共享状态。如果我们要共享,必须使用std::ref/std::cref引用包装器,线程会拷贝引用包装器,引用包装器内部的指针会指向原变量,从而实现共享。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
#include <thread>
#include <iostream>
#include <functional>
#include <string>
void append_suffix(std::string& str, const std::string& suffix)
{
str += suffix;
}
int main()
{
std::string message = "Hello";
std::string suffix = " World";
std::thread t(append_suffix, std::ref(message), std::cref(suffix));
t.join();
std::cout << message << "\n"; // 输出 "Hello World"
return 0;
}
线程的生命周期与引用对象的生命周期
有一大类并发bug的来源都是因为:线程的生命周期超过了它所引用对象的生命周期。解决方案就是复制到线程,或用shared_ptr延长生命周期。更好的方案就是在绝大多数场景下不要使用detach,使用join和RAII来配合。
线程所有权与RAII
std::thread 是不可复制的。你不能把一个线程对象赋值给另一个,也不能通过值传递来转移它。所以 std::thread 只支持 move 语义。
joining_thread:接管所有权的 RAII wrapper
wrapper 拥有线程,wrapper 析构时自动 join。这个思路的实现就是 joining_thread,它本质上就是 C++20 std::jthread
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
#include <thread>
#include <utility>
class JoiningThread {
public:
JoiningThread() noexcept = default;
// 接受任意可调用对象和参数,直接构造线程
template <typename Callable, typename... Args>
explicit JoiningThread(Callable&& func, Args&&... args)
: thread_(std::forward<Callable>(func), std::forward<Args>(args)...)
{}
// 从 std::thread move 构造——接管所有权
explicit JoiningThread(std::thread t) noexcept
: thread_(std::move(t))
{}
// 支持从另一个 JoiningThread move
JoiningThread(JoiningThread&& other) noexcept
: thread_(std::move(other.thread_))
{}
JoiningThread& operator=(JoiningThread&& other) noexcept
{
if (this != &other) {
// 先处理当前持有的线程
if (joinable()) {
join();
}
thread_ = std::move(other.thread_);
}
return *this;
}
// 也可以从一个新的 std::thread 赋值
JoiningThread& operator=(std::thread other) noexcept
{
if (joinable()) {
join();
}
thread_ = std::move(other);
return *this;
}
~JoiningThread()
{
if (joinable()) {
join();
}
}
void join()
{
thread_.join();
}
void detach()
{
thread_.detach();
}
bool joinable() const noexcept
{
return thread_.joinable();
}
// 获取底层 std::thread(用于 native_handle 等)
std::thread& get() noexcept { return thread_; }
const std::thread& get() const noexcept { return thread_; }
// 禁止复制
JoiningThread(const JoiningThread&) = delete;
JoiningThread& operator=(const JoiningThread&) = delete;
private:
std::thread thread_;
};
虽然我们实现了一个自动join的RAII 包装器,但是join()本身也是会抛出异常的,所以需要在析构函数中用try-catch把join()包起来。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
~JoiningThread()
{
if (joinable()) {
try {
join();
}
catch (const std::system_error& e) {
// join 失败了,记录日志但不抛出
// 在实际项目中应该用正式的日志系统
std::fprintf(stderr,
"JoiningThread: join() failed: %s\n", e.what());
}
}
}
C++20 的 std::jthread
C++20 标准终于引入了 std::jthread,它的行为跟我们的 JoiningThread 非常相似——析构时自动 join。但 std::jthread 还多了一个重要的功能:协作式取消(cooperative cancellation),它内部持有一个 std::stop_source,可以通过 request_stop() 请求线程停止执行。
thread_local
thread_local 就是线程存储持续时间的说明符——被它修饰的变量在每个线程中都有自己独立的实例,线程从创建到退出期间一直存在。
thread_local 变量的初始化发生在每个线程首次使用(ODR-use)时,而不是程序启动时。这个”首次使用时初始化”的行为非常重要——它保证了以下几点:第一,如果一个 thread_local 变量从来没有被某个线程访问过,那个线程就不会为它分配内存或执行初始化,所以不会有浪费。第二,初始化是线程安全的——标准保证即使多个线程同时首次访问同一个 thread_local 变量,每个线程的初始化也只会执行一次,而且不会互相干扰。
这个”延迟初始化”的特性让 thread_local 非常适合实现一些”按需分配”的资源——比如每个线程的随机数生成器、内存池、日志缓冲区等。这些资源如果全局共享就需要加锁,而用 thread_local 之后就完全无锁了,下面是一个thread_local 在性能优化中的典型用法:给每个线程一个小的内存池,小对象的分配直接从线程本地池中取,不用跟其他线程竞争
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
#include <vector>
#include <cstddef>
class ThreadLocalPool {
public:
static ThreadLocalPool& instance()
{
thread_local ThreadLocalPool pool;
return pool;
}
void* allocate(std::size_t size)
{
if (size <= kBlockSize) {
if (!free_list_.empty()) {
void* ptr = free_list_.back();
free_list_.pop_back();
return ptr;
}
// 从大块中切出一块
if (current_offset_ + size > kChunkSize) {
chunks_.emplace_back(new char[kChunkSize]);
current_offset_ = 0;
}
void* ptr = chunks_.back().get() + current_offset_;
current_offset_ += size;
return ptr;
}
// 超过块大小的分配,回退到全局分配器
return ::operator new(size);
}
void deallocate(void* ptr, std::size_t size)
{
if (size <= kBlockSize) {
free_list_.push_back(ptr);
}
else {
::operator delete(ptr);
}
}
private:
ThreadLocalPool() = default;
static constexpr std::size_t kBlockSize = 256;
static constexpr std::size_t kChunkSize = 4096;
std::vector<std::unique_ptr<char[]>> chunks_;
std::vector<void*> free_list_;
std::size_t current_offset_{kChunkSize}; // 初始值触发首次分配
};
std::call_once与std::once_flag
std::call_once 是 C++11 提供的一次性初始化机制。你给它一个 std::once_flag 和一个可调用对象,它保证无论有多少个线程同时调用 call_once,可调用对象只被执行一次——第一个到达的线程执行初始化,其余线程等待它完成。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
#include <mutex>
#include <iostream>
#include <thread>
std::once_flag init_flag;
int* shared_resource = nullptr;
void ensure_initialized()
{
std::call_once(init_flag, []() {
std::cout << "Initializing shared resource...\n";
shared_resource = new int(42);
});
}
void use_resource(const char* thread_name)
{
ensure_initialized();
std::cout << thread_name << ": resource = " << *shared_resource << "\n";
}
int main()
{
std::thread t1(use_resource, "Thread-A");
std::thread t2(use_resource, "Thread-B");
std::thread t3(use_resource, "Thread-C");
t1.join();
t2.join();
t3.join();
delete shared_resource;
return 0;
}
输出中会发现 “Initializing shared resource…” 只出现了一次——无论三个线程的调度顺序如何,初始化代码只执行了一次。std::once_flag 记录了初始化是否已经完成,call_once 在每次调用时检查这个标志位。如果初始化还没开始,第一个线程执行初始化;如果正在进行中,其他线程阻塞等待;如果已经完成,所有线程直接跳过。
std::call_once 有一个很关键的行为:如果初始化函数(可调用对象)抛出了异常,call_once 不会把 once_flag 标记为”已完成”。这意味着下一次有线程调用 call_once 时,初始化会再次尝试。
从 C++11 开始,函数作用域中的 static 局部变量有一个非常重要的保证:它的初始化是线程安全的。如果有多个线程同时首次执行到 static 变量的声明处,只有一个线程会执行初始化,其他线程会等待。