Transportmetrica B: Transport Dynamics, 13(1), 2535750
Resumen
The liberalisation of the European passenger railway markets through the European Directive EU 91/440/EEC states a new scenario where different Railway Undertakings compete with each other in a bidding process for time slots. The infrastructure resources are provided by the Infrastructure Manager, who analyses and assesses the bids received, allocating the resources to each Railway Undertaking. Time slot allocation is a fact that drastically influences the market equilibrium. In this paper, we address the time slot allocation problem within the context of a liberalized passenger railway market as a multi-objective model. The Infrastructure Manager is tasked with selecting a point from the Pareto front as the solution to the time slot allocation problem. We propose two criteria for making this selection: the first one allocates time slots to each company according to a set of priorities, while the second one introduces a criterion of fairness in the treatment of companies to incentive competition. The assessment of the impact of these rules on market equilibrium has been conducted on a liberalized high-speed corridor within the Spanish railway network.
In deregulated railway markets, efficient management of infrastructure charges is essential for sustaining railway systems. This study sets out a method for infrastructure managers to price access to railway infrastructure, focusing on freight transport in deregulated market contexts. The proposed methodology integrates negative externalities directly into the pricing structure in a novel way, balancing economic and environmental objectives. It develops a dynamic freight flow model to represent the railway system, using a logit model to capture the modal split between rail and road modes based on cost, thereby reflecting demand elasticity. The model is temporally discretized, resulting in a mesoscopic, discrete-event simulation framework, integrated into an optimization model that determines train path charges based on real-time capacity and demand. This approach aims both to maximize revenue for the infrastructure manager and to reduce the negative externalities of road transport. The methodology is demonstrated through a case study on the Mediterranean Rail Freight Corridor, showcasing the scale of access charges derived from the model. Results indicate that reducing track-access charges can yield substantial societal benefits by shifting freight demand to rail. This research provides a valuable framework for transport policy, suggesting that externality-sensitive infrastructure charges can promote more efficient and sustainable use of railway infrastructure.
The application of kernel-based Machine Learning (ML) techniques to discrete choice modelling using large datasets often faces challenges due to memory requirements and the considerable number of parameters involved in these models. This complexity hampers the efficient training of large-scale models. This paper addresses these problems of scalability by introducing the Nyström approximation for Kernel Logistic Regression (KLR) on large datasets. The study begins by presenting a theoretical analysis in which: (i) the set of KLR solutions is characterised, (ii) an upper bound to the solution of KLR with Nyström approximation is provided, and finally (iii) a specialisation of the optimisation algorithms to Nyström KLR is described. After this, the Nyström KLR is computationally validated. Four landmark selection methods are tested, including basic uniform sampling, a k-means sampling strategy, and two non-uniform methods grounded in leverage scores. The performance of these strategies is evaluated using large-scale transport mode choice datasets and is compared with traditional methods such as Multinomial Logit (MNL) and contemporary ML techniques. The study also assesses the efficiency of various optimisation techniques for the proposed Nyström KLR model. The performance of gradient descent, Momentum, Adam, and L-BFGS-B optimisation methods is examined on these datasets. Among these strategies, the k-means Nyström KLR approach emerges as a successful solution for applying KLR to large datasets, particularly when combined with the L-BFGS-B and Adam optimisation methods. The results highlight the ability of this strategy to handle datasets exceeding 200,000 observations while maintaining robust performance.
This paper presents a software package developed in Python that allows the application of the technique known as Kernel Logistic Regression (KLR), a Machine Learning (ML) tool, to the problem of transport demand prediction. More concretely, it permits the specification of a series of models using KLR and their estimation by means of a Penalised Maximum Likelihood Estimation (PMLE) procedure providing a set of goodness-of-fit indicators and the application of model validation techniques. Another functionality is that it allows to extract from the model several indicators such as the Willingness to Pay (WTP) or the Value of Time (VOT).
The use of simulators is frequent in the study of complex systems. They replicate a real system and allow obtaining data from the simulated process as well as providing a mathematical model that helps answer What If? questions. This opens up the possibility of evaluating laboratory environment techniques and processes that will later be implemented in the real system. Competitive passenger rail services are complex systems. The efficient design of market mechanisms that encourage competition to offer more and better services requires analytical tools on which to test these new policies. This makes it essential to develop a simulator, as real and complete as possible, to carry out this analysis. This paper presents the development of a microscopic simulator, called ROBIN (Rail mOBIlity simulatioN), to simulate rail mobility in a competitive regime. The developed simulator is parameterizable and general enough to simulate passenger flows in competing railway systems. The tests carried out with ROBIN using the Spanish railway market as a case of study are proof of its usefulness, allowing disaggregated modelling of passenger behaviour and their travel choices in a competitive railway market.
Transportation Research Part C: Emerging Technologies, 156, 104318
Resumen
The emergence of a variety of Machine Learning (ML) approaches for travel mode choice prediction poses an interesting question to transport modellers: which models should be used for which applications? The answer to this question goes beyond simple predictive performance, and is instead a balance of many factors, including behavioural interpretability and explainability, computational complexity, and data efficiency. There is a growing body of research which attempts to compare the predictive performance of different ML classifiers with classical Random Utility Models (RUMs). However, existing studies typically analyse only the disaggregate predictive performance, ignoring other aspects affecting model choice. Furthermore, many existing studies are affected by technical limitations, such as the use of inappropriate validation schemes, incorrect sampling for hierarchical data, a lack of external validation, and the exclusive use of discrete metrics. In this paper, we address these limitations by conducting a systematic comparison of different modelling approaches, across multiple modelling problems, in terms of the key factors likely to affect model choice (out-of-sample predictive performance, accuracy of predicted market shares, extraction of behavioural indicators, feature importance analysis, and computational efficiency). The modelling problems combine several real world datasets with synthetic datasets, where the data generation function is known. The results indicate that the models with the highest disaggregate predictive performance (namely Extreme Gradient Boosting (XGBoost) and Random Forests (RF)) provide poorer estimates of behavioural indicators and aggregate mode shares, and are more expensive to estimate, than other models, including Deep Neural Networks (DNNs) and Multinomial Logit (MNL). It is further observed that the MNL model performs robustly in a variety of situations, though ML techniques can improve the estimates of behavioural indices such as Willingness To Pay (WTP).
Transport demand modelling plays a critical role in transportation planning, enabling the accurate prediction of future transport demand and the evaluation of transport policies and infrastructure plans. However, as transport systems become increasingly complex and recent advances in technology result in massive data collection, traditional analytical methods like Random Utility Models (RUMs) are no longer sufficient to manage this complexity. Therefore, it is necessary to incorporate new techniques to overcome this limitation. This thesis investigates the potential of Machine Learning (ML) methods in this context. Firstly, it is analysed whether state-of-the-art ML models such as artificial neural networks, support vector machines, and ensemble methods like random forests or gradient boosting decision trees, are superior to RUMs in this research field. To achieve this, the models are compared considering as differential criteria the predictive performance and the ability to derive indicators of decision-makers’ behaviour, always in the context of transport demand modelling. The results show that classical techniques are outperformed by ML models, but also show that the latter have difficulties in generating reliable econometric indicators. For this reason, a ML model called Kernel Logistic Regression (KLR) is proposed as an alternative to model the utility functions of RUMs, enabling the derivation of econometric indicators. The experiments conducted demonstrate that KLR provides good results on real-world datasets used in previous comparisons in the literature, while providing unbiased estimates of behavioural indicators. Additionally, it is proposed to extend the application of KLR to a wider range of ML problems by means of an extension of the KLR models called Generalized Kernel Logistic Regression (GKLR). For instance, the GKLR theory has led to the derivation of a novel model called Nested Kernel Logistic Regression (NKLR), which enables the application of KLR to datasets with hierarchically structured data. Finally, this thesis addresses one of the main limitations of the KLR method, which is the high computational and spatial complexity in large-scale problems. To overcome this limitation, it is suggested the use of the Nyström technique and the implementation of accelerated versions of line search training methods. The results demonstrate that by incorporating these techniques, KLR can efficiently tackle large-scale problems involving hundreds of thousands of data points.
Traditionally, Random Utility Maximization (RUM) models have been widely applied to travel mode choice modelling. Currently, Machine Learning (ML) models are being applied as an alternative to RUM models, since they provide better results in terms of prediction capability and they can manage large volumes of data. In this paper, a comprehensive comparison between classic RUM models and ML models, including single and ensemble classifiers as well as Deep Neural Networks (DNNs), is provided in order to assess systematically the performance of different models over two different datasets which have different sizes and nature of data. Numerical experiments show Random Forest (RF) is the best classifier in terms of accuracy index and the computational cost to train the model.
Dynamic traffic management (DTM) systems are used to reduce the negative externalities of traffic congestion, such as air pollution in urban areas. They require traffic and environmental monitoring infrastructures. In this paper we present a prototype of a low-cost Internet of Things (IoT) system for monitoring traffic flow and the Air Quality Index (AQI). The computation of the traffic flows is based on processing video in the compressed domain. Only using motion vectors as input, traffic flow is computed in real-time over an embedded architecture. An estimation of the AQI is supported by machine learning regression techniques, using different feature data obtained from the IoT device. These automatic learning techniques overcome the need for complex calibration and other limitations of embedded devices in making the needed measurements of the pollutant gases for the computation of the actual AQI. The experimentation with the data obtained from different cities representing different scenarios with a variety of climate and traffic conditions, allows validating the proposed architecture. As regressors, Linear Regression (LR), Gaussian Process Regression (GPR) and Random Forest (RF) are compared using the performance metrics R2, MSE, MAE and MRE resulting in a relevant improvement of the AQI estimations of our proposal.
In the last few years, the success of Machine Learning (ML) algorithms has led to the extension of their applications to areas such as transport planning. One of the main tasks within transport planning is the analysis of transport demand. To do so, it is necessary to analyse the way in which users make their decisions about the trips they make and, therefore, be able to predict the number of passengers on the transport network in relation to respect to interventions made on the transport system. Consequently, transport policies and plans can be evaluated according to the behaviour of the passengers. Discrete choice models based on random utility maximization have been developed over the last four decades, becoming the canonical tool for transport demand analysis. Nowadays, the use of ML methods could provide an alternative to discrete choice models, as they reduce the need for the analyst to specify the functional expression of these models and achieve a higher level of accuracy in their predictions. A Python software package called PyKernelLogit was developed to apply a ML method called Kernel Logistic Regression (KLR) to the problem of predicting the transport demand. This package allows the user to specify a set of models using KLR and the estimation of those using a Penalized Maximum Likelihood Estimation procedure. Moreover, this tool also provides a set of indicators for goodness of fit and the application of model validation techniques. Finally, it allows to obtain the willingness to pay or value of time indicators commonly used in transport planning.
The success of machine-learning methods is spreading their use to many different fields. This paper analyses one of these methods, the Kernel Logistic Regression (KLR), from the point of view of Random Utility Model (RUM) and proposes the use of the KLR to specify the utilities in RUM, freeing the modeler from the need to postulate a functional relation between the features. A Monte Carlo simulation study is conducted to empirically compare KLR with the Multinomial Logit (MNL) method, the Support Vector Machine (SVM) and the Random Forests (RF). We have shown that, using simulated data, KLR is the only method that achieves maximum accuracy and leads to an unbiased willingness-to-pay estimator for non-linear phenomena. In a real travel mode choice problem, RF achieved the highest predictive accuracy, followed by KLR. However, KLR allows for the calculation of indicators such as the value of time, which is of great importance in the context of transportation.
Computational Statistics & Data Analysis, 144, 106844
Resumen
Several common general purpose optimization algorithms are compared for finding A- and D-optimal designs for different types of statistical models of varying complexity, including high dimensional models with five and more factors. The algorithms of interest include exact methods, such as the interior point method, the Nelder–Mead method, the active set method, the sequential quadratic programming, and metaheuristic algorithms, such as particle swarm optimization, simulated annealing and genetic algorithms. Several simulations are performed, which provide general recommendations on the utility and performance of each method, including hybridized versions of metaheuristic algorithms for finding optimal experimental designs. A key result is that general-purpose optimization algorithms, both exact methods and metaheuristic algorithms, perform well for finding optimal approximate experimental designs.
The Kernel Logistic Regression is a popular technique in machine learning. In this work this technique is applied to the field of discrete choice modeling. This approach is equivalent to specifying non-parametric utilities in random utility models. A Monte Carlo simulation experiment has been carried out to compare this approach with Multinomial Logit models, comparing the goodness of fit and the capability of obtaining the specified utilities.
2019 28th International Conference on Computer Communication and Networks (ICCCN), 1-8
Resumen
The present work addresses the problem of community detection in social networks. Determination of the community structure permits to identify the organizational and functional units of the network. To such an end, the most powerful approach is the fuzzy one. Here, each network node participates in all the communities simultaneously. In this paper, we develop a general formalism for fuzzy community detection in weighted, unweighted, directed and undirected social networks. Using an error functional measuring the difference between the predicted and observed network topologies, the problem is transformed into an unconstrained minimization. This allows to define a general algorithmic pattern based on the greedy paradigm, which can be customized to several different fuzzy community detection algorithms. Analysis of the algorithmic complexity of the procedure permits to determine the factors responsible for the variation of its running time. Using this information, we define an efficient algorithm (TRIBUNE) by applying a conjugate gradient approach. To determine the performance of TRIBUNE, we introduce a new network model specially designed as a fuzzy community detection benchmark. Using this benchmark, we show that TRIBUNE greatly reduces the dependence on the number of iterations of the optimization procedure. As a result, TRIBUNE needs in average 84% less time than the steepest descent baseline to solve a test set of benchmark networks. Furthermore, the performance of TRIBUNE increases with the network size.
⌕
No se encontraron publicaciones
Prueba con otra búsqueda o elimina alguno de los filtros.