Data Science
Fine-Tuning Language Models for Sentiment Analysis in Financial Markets
13 minutes
In markets, as in life, many decisions described as rational are actually driven by hidden psychological and behavioral factors. Recognizing how emotions affect decision-making is one of the things that separates good investors from excellent ones.
Behavioral finance is a field that studies how psychological and behavioral factors affect financial decisions. It sits at the intersection of finance, psychology, and economics, and seeks to understand how investors make decisions and how those decisions affect markets.
Even though global financial markets are increasingly dominated by algorithms and systematic investment strategies, behavioral biases remain very much present. These are patterns of behavior that can lead investors to make irrational or suboptimal decisions. Examples include confirmation bias, herd behavior, anchoring, and hindsight bias.
These biases are common among individual investors, but can also affect institutional and professional investors. Financial markets are influenced by external factors such as news, economic events, and political changes, which can amplify them.
In this article, I'll explore how language models can be used to predict financial-market sentiment from monthly or quarterly letters written by asset managers. These letters are a rich source of information about professionals' views and can offer valuable insight into market sentiment.
I'll cover language models and transformers for text classification and sentiment analysis, transfer learning and fine-tuning for domain adaptation, and PEFT (Parameter-Efficient Fine-Tuning) for training models on specific tasks.
The complete project and code are available on GitHub.
Methodology
There are several ways to measure financial-market sentiment. One common approach is to use quantitative indicators: numerical metrics intended to capture the market's temperature. A well-known example is CNN's Fear & Greed Index, which measures the sentiment of US investors on a scale from 0 (fear) to 100 (greed).
The index combines seven indicators, including volatility, moving averages, the volume of advancing and declining stocks, put and call options, and the number of stocks reaching highs and lows.
When I wrote the original article, the index stood at 21 (extreme fear), amid the Trump administration's decisions to impose new tariffs on products from other countries, especially China.
Other approaches use qualitative data such as news, financial reports, and social media. Much of this information is unstructured, so it requires specific techniques from natural language processing (NLP) and artificial intelligence.
This project was strongly inspired by Lucas Leme's paper, which used news from financial outlets such as Valor Econômico and Infomoney to train a model to predict market sentiment.
Source: FinBERT-PT-BR: Sentiment Analysis of Portuguese Financial Market Texts
Financial news is a rich source of information and produces a large volume of data, but it expresses the media's view of the market, not necessarily investors' views. For this project, we'll use monthly and quarterly letters from independent asset managers to build a dataset for a market-sentiment indicator.
Asset-management firms are specialized institutions that professionally manage third-party financial assets. They raise money from investors and allocate it across asset classes such as equities, fixed income, real estate, and other investments. Analysts and portfolio managers decide how to invest those resources according to a defined strategy.
A common practice is to publish periodic letters to investors, usually monthly or quarterly, describing fund performance, strategies, and future outlook. These letters offer a valuable perspective on the views of important market participants and are closely followed by companies and academics.
The first step was to collect these letters directly from the managers' websites. I used BeautifulSoup to scrape the sites and extract text from PDFs. This was a demanding step: letters came in different formats and layouts, requiring data cleaning and normalization, and each firm's website required its own scraper adjustments.
In total, I collected 707 letters from 12 firms, covering 1999–2025. Firms were selected based on their relevance in the Brazilian market and the availability of letters online. I focused on independent managers, unaffiliated with financial institutions, with a stronger equity orientation.
| Asset manager | Letters |
|---|---|
| Guepardo | 111 |
| IP Capital | 95 |
| Dahlia Capital | 81 |
| Dynamo | 80 |
| Kapitalo | 74 |
| Ártica | 62 |
| Encore | 58 |
| Genoa Capital | 56 |
| Alpha Key | 35 |
| Mar Asset | 20 |
| Alaska | 18 |
| Squadra | 17 |
| Total | 707 |
The extracted text was cleaned and normalized, with irrelevant material such as tables, charts, and images removed. I stored the text in a SQLite database along with metadata such as title and date.
Transformers and language models
Language models are deep-learning models designed to understand and generate text. They are based on the transformer architecture, introduced in the 2017 paper "Attention Is All You Need".
I won't explain transformer architecture in detail; there are many accessible articles and videos on the subject. I recommend Andrej Karpathy's YouTube channel, DeepLearning.AI's Generative AI with LLMs course, and the Hugging Face LLM course.
At a high level, these models are large neural networks trained on massive text collections. Given a sequence of words, they can predict the next word. Not every model works this way, but it's a useful introduction because GPT, the best-known example, does.

There are three main types of language models: encoder-only, decoder-only, and encoder-decoder. GPT (used by ChatGPT) and Llama are decoder-only models specialized in generating text from a prompt. They are trained with causal language modeling: given a sequence, predict the next word. In this sense, they process text in one direction.
Encoder-only models such as BERT are designed to understand a word in context, considering both the words that come before and those that follow. This lets them better capture meaning and relationships in text.
BERT is trained using masked language modeling: some words are masked, and the model must predict them from the surrounding words. This gives it a bidirectional view of text, which is useful for tasks such as text classification, question answering, and sentiment analysis.

That's the kind of model we'll use—or, more precisely, a version called BERTimbau, a Portuguese model based on Google's BERT. Developed by NeuralMind AI, BERTimbau is available in base (110M) and large (335M) parameter versions. In general, more parameters mean a more capable model.
BERTimbau cannot predict financial-market sentiment out of the box. We need to train it for text classification, and it was not specifically trained on financial material. It is also not an LLM with billions of parameters, which may limit its ability to understand the context and meaning of words in this domain.
We'll therefore use BERTimbau as a base model and apply transfer learning in two stages:
- Domain-Adaptive Pretraining (DAPT): Continue pretraining the model on a large collection of financial text so it can learn domain-specific vocabulary and context. This stage uses a general objective such as word prediction.
- Task fine-tuning: Adapt the model to a specific task—in this case, text classification and sentiment prediction—using labeled examples.
Domain-adaptive pretraining
The goal of this stage is to teach the model what it needs to know about finance and investing. Before asking a person to classify the sentiment of a text, we'd want them to have a solid grounding in the subject. The same is true of a language model.
To understand financial text, a model needs to know domain-specific terms and jargon such as fixed income, interest rates, stocks, the Central Bank, volatility, alpha, and beta. It also needs to understand how these words are used and related to each other.
Our base model has only 110M parameters, so we shouldn't expect deep knowledge of specialized domains. Given the phrase “bull market,” for example, the model might associate it with bulls rather than a rising market.
The example is amusing, but domain adaptation matters for sentiment classification. We continued the original BERT masked-language-modeling task, this time using excerpts from the asset managers' letters. We used 36,000 excerpts, averaging about 200 characters each.
The code is available in the project repository. I trained the model on an AWS g4dn.xlarge GPU instance using Hugging Face's Transformers library. Training took about two hours.
The adapted model's predictions are revealing:
from transformers import AutoModelForMaskedLM, AutoTokenizer
from transformers import pipeline
base_model = AutoModelForMaskedLM.from_pretrained('neuralmind/bert-base-portuguese-cased')
domain_adapted_model = AutoModelForMaskedLM.from_pretrained('../models/bert-portuguese-asset-management')
tokenizer = AutoTokenizer.from_pretrained('neuralmind/bert-base-portuguese-cased', do_lower_case=False)
def predict_mask(text: str, model: str, top_k: int = 5):
fill_mask = pipeline("fill-mask", model=model, tokenizer=tokenizer, top_k=top_k)
outputs = fill_mask(text)
for o in outputs:
token = o["token_str"].strip()
score = o["score"]
print(f" {token:15s} → {score:.4f}")
prompt = "Houve um debate interno sobre alocar mais recursos em [MASK]."
print("======= Base Model =======\n")
predict_mask(prompt, base_model, top_k=5)
# educação 0.0856; projetos 0.0781; saúde 0.0671; hospitais 0.0483; infraestrutura 0.0477
print("\n=== Domain Adapted Model ===\n")
predict_mask(prompt, domain_adapted_model, top_k=5)
# ações 0.3805; empresas 0.0754; juros 0.0431; investimentos 0.0409; infraestrutura 0.0288
The adapted model has a much stronger grasp of the financial context. The base model completes the sentence with words such as education, health, and hospitals; the adapted model suggests stocks, companies, and interest rates. This illustrates the impact of DAPT. We can also evaluate the model using metrics such as perplexity.
Perplexity is the exponential of cross-entropy (or loss) on a test set. Lower perplexity means the model is less uncertain about which token should fill the [MASK].
| Base Model (BERTimbau) | Domain-Adapted Model | |
|---|---|---|
| Perplexity | 10.95 | 4.29 |
On unseen test data, the base model's perplexity was 10.95, compared with 4.29 for the adapted model. The result suggests that adaptation helped the model predict financial-domain text more accurately.
PEFT and task-specific fine-tuning
The second stage teaches the model to classify sentiment. We'll take the domain-adapted model and fine-tune it for text classification.
So far, the model can represent context and reconstruct masked tokens. To classify text, we add a classification head that maps the model's output to the sentiment classes. In simple terms, a linear layer produces logits—unnormalized values for each class—and a softmax turns them into probabilities. The class with the highest probability becomes the prediction:
[
{"label": "NEGATIVE", "score": 0.85},
{"label": "NEUTRAL", "score": 0.10},
{"label": "POSITIVE", "score": 0.05}
]
Training requires labeled examples of positive, negative, and neutral text. I normalized the collected letters, split them into excerpts averaging 200 characters, randomly sampled 1,000 examples without stratification, and manually labeled each one as POSITIVE, NEGATIVE, or NEUTRAL.
A more rigorous process would require several safeguards:
- Stratified sampling: Ensure each relevant group is represented. Some managers have many more letters than others, so an unbalanced sample could make the model overly familiar with a few firms' writing styles.
- Class balancing: Ensure each sentiment class has enough examples. If most excerpts are neutral—because they contain factual or technical information—the model may learn to over-predict that class.
- Multiple annotators: Manual labeling is subjective. Multiple annotators and an agreement measure such as Cohen's kappa would help assess consistency, ideally with input from domain professionals.
I kept the process simple because the goal was to demonstrate fine-tuning, not to build a production model. The manual labels inevitably reflect human judgment and do not necessarily represent market sentiment objectively. With more careful data collection and annotation, this approach could be developed into a useful tool for analysts and asset managers.
Here are examples of positive and negative excerpts, translated from the original Portuguese dataset:
POSITIVE
- “Although we have no control over short-term returns, we are very pleased with our investments—regardless of which way politics turns.”
- “It's hard to remember a time when we could build equity portfolios with such an attractive return-and-risk profile.”
- “Although the new law has not yet been approved, the discussion around it has already produced positive effects that could benefit Brazil's capital markets.”
NEGATIVE
- “Through imitation, in the most traditional herd behavior, selling at any price now seems like the rational thing to do.”
- “The current environment is challenging, especially for investors who need a continuous income stream, such as retirees and pension funds.”
- “Above all, although we are very cautious and concerned about the market, we are convinced that we are doing the right things.”
- “There is still considerable doubt about the new government's ability to deliver fiscal adjustment and how quickly it can do so.”
Positive excerpts generally express optimism, confidence, and hope about the market and economy. Negative ones express pessimism, distrust, fear, and uncertainty.
After creating the labeled dataset, I trained the model using PEFT (Parameter-Efficient Fine-Tuning), specifically LoRA (Low-Rank Adaptation). Fine-tuning can require substantial memory: in addition to model parameters, training stores gradients, optimizer states, and other intermediate values. This can overwhelm GPUs with limited memory.
LoRA freezes most model parameters, making it possible to fine-tune with less memory and compute. PEFT methods operate on selected layers, so the model can adapt to a task without updating every parameter. Our model is not especially large, but LoRA still reduces training time and memory use while preserving much of its performance.
The output is not a new standalone model but an adapter containing the parameters learned for the task. It is loaded alongside the unchanged base model. The same base model can be reused for other tasks with different adapters, which are much smaller than the base model.

I trained the model on a remote GPU instance. The code is in the project repository. Using Hugging Face's PEFT library, training updated only 1,181,955 parameters—just over 1% of the base model's parameters.
The trained adapter can be combined with the base model to create the final classifier. The model is available on my Hugging Face Hub profile.
Model evaluation
I compared our model (Model AM) with a classifier built from BERTimbau using the same data and a held-out test set. The goal was to assess whether domain-adaptive pretraining improved performance over the base model. I used accuracy and F1 score. I also included FinBERT-PT-BR, introduced earlier, a BERT-based financial sentiment model.
| Model | Accuracy | F1 score |
|---|---|---|
| Model AM | 0.7300 | 0.4582 |
| FinBERT-PT-BR | 0.4800 | 0.3401 |
| BERTimbau | 0.3200 | 0.2249 |
Model AM substantially outperformed BERTimbau, which had not been adapted to financial text. FinBERT-PT-BR performed worse than our model, perhaps because it was trained on a different kind of data (news) and does not capture the same domain.
The goal was to predict sentiment in asset managers' letters, not sentiment in the broader market or media. Although our model outperformed the alternatives, it is not yet ideal: 73% accuracy and an F1 score of 45% leave plenty of room for improvement.
An obvious next step would be to increase the amount of labeled data and training time. Expanding the dataset to 5,000 or 10,000 examples, along with input from additional annotators, could lead to a substantial improvement.
A less obvious issue is class imbalance. In our training set, 67% of examples were labeled NEUTRAL. Much of an asset-management letter is indeed neutral, but the model can learn to favor that class. Undersampling could balance the classes, but would sharply reduce the number of labeled examples. Instead, I used class weighting, assigning more weight to errors on the less represented POSITIVE and NEGATIVE classes:
class WeightedTrainer(Trainer):
def compute_loss(self, model, inputs, num_items_in_batch=None, return_outputs=False):
labels = inputs.get("labels")
outputs = model(**inputs)
logits = outputs.get("logits")
class_weights = torch.tensor([2.0, 2.0, 1.0], device=logits.device)
loss_fct = torch.nn.CrossEntropyLoss(weight=class_weights)
loss = loss_fct(logits.view(-1, self.model.config.num_labels), labels.view(-1))
return (loss, outputs) if return_outputs else loss
During experimentation, adjusting the loss function helped improve the model, especially its F1 score, which is more informative for imbalanced classification. Combined with more labeled data and annotators, these techniques could lead to a much stronger model.
We can now use the model for inference—predicting the sentiment of new text—and, for example, build a financial-market sentiment indicator from asset managers' letters dating back to 1999.

I assigned a sentiment score to each letter and calculated an average for each quarter, which is usually the frequency at which letters are published alongside companies' financial results. The chart shows a 12-month moving average on a scale from 0 (very negative) to 100 (very positive). The series has large fluctuations, especially in the early 2000s, when few letters were available.
To map the model's probabilities to a score between 0 and 100, I removed the neutral probability and normalized the positive and negative probabilities. The score for each letter is calculated as follows:
def _aggregate(prob_chunks: List[List[Dict[str, float]]]) -> int:
if not prob_chunks:
return 50
scores, weights = [], []
for chunk in prob_chunks:
d = {x["label"]: x["score"] for x in chunk}
p_pos, p_neg = d["POSITIVE"], d["NEGATIVE"]
s_i = p_pos - p_neg
w_i = p_pos + p_neg
if w_i:
p_pos, p_neg = p_pos / w_i, p_neg / w_i
s_i = p_pos - p_neg
scores.append(s_i)
weights.append(w_i)
S = np.dot(scores, weights) / sum(weights) if weights else 0.0
S = np.sign(S) * abs(S) ** 0.75 # alpha = 0.75
return int(50 * (S + 1)) # [-1,1] -> [0,100]
Despite the fluctuations, the result shows interesting patterns and suggests a promising methodology with plenty of room for improvement.
Conclusion
This project shows that even in an increasingly quantitative market, emotions and human behavior still matter—and that the right tools may help us measure them.
We adapted a language model to financial vocabulary and then taught it to classify asset managers' letters by sentiment. Despite limitations in labeled data and class imbalance, the model produced promising results. We used it to build a market-sentiment indicator spanning more than two decades, revealing interesting patterns and opening the door to further work.
The experiment illustrates how emerging AI techniques can create tangible links between technology and financial markets. I trained the models on an AWS g4dn.xlarge instance with 16 GB of GPU memory and four GPU cores. It came with common deep-learning frameworks such as TensorFlow and PyTorch and NVIDIA drivers, but I used Docker to make the environment reproducible.
Including time spent learning and experimenting, the total cost was 108.84). Overall, this was a valuable project that taught me a lot about language models, transfer learning, and the Hugging Face ecosystem. I hope it was useful to you as well.
Thanks for reading. If you have any questions or suggestions, feel free to contact me by email, LinkedIn or Twitter.
Let's talk
Have a process that could work better, a data question or an application idea? Share the context on LinkedIn. I enjoy exchanging ideas about software, data and AI challenges.
Connect with me on LinkedIn